EADST

QWEN7B to LLAMA GPTQ model structure

Here is the markdown format for the GPTQ model structure, detailing each layer and component:


GPTQ Model Structure

The GPTQ model consists of the following layers and components:

Embedding Layer

  • model.embed_tokens.weight: torch.Size([151851, 4096])

Layers

Each layer in the model has the following components:

Layer 0 to Layer 31

Each layer (model.layers.[0-31]) includes:

  • input_layernorm.weight: torch.Size([4096])

  • Self-Attention Sublayer:

    • k_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

    • o_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

    • q_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

    • v_proj:

      • qweight: torch.Size([512, 4096])

      • qzeros: torch.Size([32, 512])

      • scales: torch.Size([32, 4096])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([4096])

  • MLP (Multi-Layer Perceptron) Sublayer:

    • down_proj:

      • qweight: torch.Size([1376, 4096])

      • qzeros: torch.Size([86, 512])

      • scales: torch.Size([86, 4096])

      • g_idx: torch.Size([11008])

      • bias: torch.Size([4096])

    • gate_proj:

      • qweight: torch.Size([512, 11008])

      • qzeros: torch.Size([32, 1376])

      • scales: torch.Size([32, 11008])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([11008])

    • up_proj:

      • qweight: torch.Size([512, 11008])

      • qzeros: torch.Size([32, 1376])

      • scales: torch.Size([32, 11008])

      • g_idx: torch.Size([4096])

      • bias: torch.Size([11008])

  • post_attention_layernorm.weight: torch.Size([4096])

Final Layer Normalization and Output

  • model.norm.weight: torch.Size([4096])
  • lm_head.weight: torch.Size([151851, 4096])
相关标签
About Me
XD
Goals determine what you are going to be.
Category
标签云
Pandas Random LaTeX GGML CLAP Magnet Vmess 继承 LeetCode FastAPI printf v2ray RAR Plotly CC Land Qwen2 Agent EXCEL CUDA Mixtral Bipartite COCO Vim Qwen Python Claude Template 论文速读 PDB FP64 FP16 顶会 VGG-16 Excel diffusers ChatGPT logger 多进程 Animate LLM 签证 TensorFlow Proxy HuggingFace Color Rebuttal Cloudreve GoogLeNet Math AI Qwen2.5 YOLO 证件照 BeautifulSoup Jupyter API FlashAttention SPIE Card IndexTTS2 mmap CEIR uWSGI Michelin Diagram Heatmap RL UI PyCharm Permission UNIX Freesound WAN CV NameSilo Markdown transformers VSCode OpenCV v0.dev 域名 关于博主 Data 递归学习法 Windows ModelScope LLAMA Bitcoin Gemma XGBoost Bert Interview Tiktoken 搞笑 Datetime PIP Domain GPT4 音频 Website 报税 Hungarian CSV Llama 论文 净利润 Pillow Numpy SAM git Translation git-lfs uwsgi News HaggingFace Git Breakpoint Plate Hilton Transformers DeepStream Nginx 云服务器 Zip C++ Sklearn JSON 图标 hf Firewall VPN Dataset Distillation WebCrawler Video LoRA Ubuntu icon InvalidArgumentError RGB Streamlit Github ms-swift Algorithm Tensor SQLite Crawler Pytorch Statistics XML llama.cpp Baidu Quantization PyTorch Image2Text QWEN TSV Hotel Password Input SVR NLTK Google OpenAI Safetensors CTC 强化学习 torchinfo BTC Disk TTS Django Clash scipy Pickle Jetson Conda 财报 Attention CAM Miniforge Logo DeepSeek OCR Food 阿里云 Knowledge Base64 NLP ResNet-50 Web tqdm Anaconda Paper Paddle 多线程 Use Bin PDF Augmentation 第一性原理 Ptyhon ONNX Quantize 图形思考法 Review 飞书 公式 MD5 SQL GPTQ Tracking tar Docker TensorRT Shortcut Search FP8 GIT 腾讯云 BF16 算法题 FP32 Linux 版权
站点统计

本站现有博文333篇,共被浏览922187

本站已经建立2628天!

热门文章
文章归档
回到顶部