EADST

QWEN7B to LLAMA GPTQ model structure

Here is the markdown format for the GPTQ model structure, detailing each layer and component:


GPTQ Model Structure

The GPTQ model consists of the following layers and components:

Embedding Layer

  • model.embed_tokens.weight: torch.Size([151851, 4096])

Layers

Each layer in the model has the following components:

Layer 0 to Layer 31

Each layer (model.layers.[0-31]) includes:

  • input_layernorm.weight: torch.Size([4096])

  • Self-Attention Sublayer:

    • k_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

    • o_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

    • q_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

    • v_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

  • MLP (Multi-Layer Perceptron) Sublayer:

    • down_proj:

      • qweight: torch.Size([1376, 4096])

      • qzeros: torch.Size([86, 512])

      • scales: torch.Size([86, 4096])

      • g_idx: torch.Size([11008])

      • bias: torch.Size([4096])

    • gate_proj:

      • qweight: torch.Size([512, 11008])

      • qzeros: torch.Size([32, 1376])

      • scales: torch.Size([32, 11008])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([11008])

    • up_proj:

      • qweight: torch.Size([512, 11008])

      • qzeros: torch.Size([32, 1376])

      • scales: torch.Size([32, 11008])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([11008])

  • post_attention_layernorm.weight: torch.Size([4096])

Final Layer Normalization and Output

  • model.norm.weight: torch.Size([4096])
  • lm_head.weight: torch.Size([151851, 4096])
相关标签
About Me
XD
Goals determine what you are going to be.
Category
标签云
Ubuntu Video SAM transformers TSV Clash TensorFlow CTC Diagram Random Color PDF PyTorch FP16 BTC Agent 第一性原理 Image2Text 公式 递归学习法 Claude 继承 Conda Hilton Qwen Vmess Transformers Jupyter SQL Bipartite 强化学习 AI IndexTTS2 CC ResNet-50 版权 git-lfs C++ tar PyCharm XGBoost GGML LeetCode 音频 Food v0.dev Pickle CUDA Augmentation OpenAI Miniforge ModelScope Firewall 报税 Google Python 论文速读 TTS CSV FP8 Heatmap Jev GPTQ Website Translation Quantization Vim icon Input Cloudreve mmap 域名 Breakpoint FP32 云服务器 VSCode 腾讯云 NameSilo Linux Bitcoin NLP Anaconda Disk CV 净利润 Numpy QWEN Use Ptyhon Gemma Github Search Card ONNX Freesound RGB Rebuttal Bin Proxy Pytorch CEIR BeautifulSoup WebCrawler XML Hungarian Jetson 多进程 Tiktoken Plotly Hotel 签证 Datetime Excel GIT Tensor GPT4 Password Git Web Tracking Data logger News Domain Safetensors Statistics uwsgi Land llama.cpp FlashAttention Plate Animate Base64 hf Logo Baidu SVR HaggingFace API网关 CAM UNIX git Nginx GoogLeNet uWSGI API RL OCR LaTeX Qwen2 多线程 PDB Zip RAR Streamlit 图标 顶会 Django FastAPI MD5 YOLO OpenCV Bert Quantize Attention HuggingFace ms-swift 证件照 v2ray 算法题 Knowledge BF16 torchinfo Markdown 搞笑 VPN FP64 LoRA UI LLM LLAMA CLAP COCO Michelin Template DeepSeek Paper SQLite Sklearn Interview Math SPIE Windows PIP tqdm 阿里云 VGG-16 TensorRT EXCEL Permission Algorithm Pandas Paddle ChatGPT WAN NLTK 关于博主 Shortcut Mixtral Docker scipy Qwen2.5 Harness 论文 DeepStream Dataset Distillation Crawler Review 图形思考法 Magnet diffusers printf InvalidArgumentError Llama Pillow 飞书 JSON 财报
站点统计

本站现有博文337篇,共被浏览961402次

本站已经建立2673天!

热门文章
文章归档
回到顶部