Sakura EmbeddingGemma 2 — AutoRound V2 HQ GGUF

⭐ Choose your HQ quality tier

All three are custom native AutoRound mixed-precision text-embedding models, not full Q8 models or unmodified S/M presets. The filename identifies the base storage family; selected PLE tensors use Q8_0. Q5/Q6 also use a Q8_0 output embedding head.

Base model: google/embeddinggemma-2, pinned snapshot 914f7f89142e33e77833254d9c9b90c3cef7303b. Architecture gemma-embedding2, 413 tensors, approximately 271M text-path parameters. Native embeddings 768d; MRL 512/256/128 after truncation and L2 renormalization.

File Details

Tier File Bytes MiB SHA256
Recommended balance: Q5 HQ sakura-embeddinggemma-2-autoround-v2-Q5_K_S-HQ-cal256.gguf 209,585,440 199.88 d0b5f556ed4916425a3512e4af1f0d7e631f9f85b6224b6b6d956c8f33adfaa0
Highest measured fidelity: Q6 HQ sakura-embeddinggemma-2-autoround-v2-Q6_K-HQ-cal256.gguf 243,008,800 231.75 e5c8680e1562e47ac5d379a25dc98972ccbd037df858dbca387c1df316ae0430
Compact: Q4 HQ sakura-embeddinggemma-2-autoround-v2-Q4_K_S-HQ-cal256.gguf 178,033,120 169.79 e01262207bf83d43043f9c880d4388b8b35f002c3fdaf86101252b3356f62ab1

Side-by-Side Comparison — same BF16 benchmark, 768d

Mean Cosine and Spearman combine 300 vectors: 100 text queries, 100 text documents, 50 code queries and 50 code snippets. Text and code ranking metrics are shown separately. All rows use the same existing reference, prefixes, metric functions and runtime flags.

Variant MiB Mean Cosine Min Cosine Spearman Text Top-5 Code Top-5 Text/Code Recall@5 Text nDCG@10 Code nDCG@10
Sakura Q4 HQ 169.79 0.99400270 0.98141599 0.98857239 90.00% 87.60% 100% / 100% 0.97668917 0.97540851
Unsloth UD-Q4_K_XL 167.54 0.99354154 0.97681332 0.98770997 89.60% 88.40% 100% / 100% 0.97352332 0.97615687
Sakura Q5 HQ 199.88 0.99828082 0.99384677 0.99664205 95.40% 92.80% 100% / 100% 0.98892991 0.99019274
Unsloth UD-Q5_K_XL 200.32 0.99833661 0.99317700 0.99675520 94.80% 94.40% 100% / 100% 0.98853356 0.99009068
Sakura Q6 HQ 231.75 0.99940187 0.99797606 0.99879852 96.80% 97.20% 100% / 100% 0.99555133 0.99365559
Unsloth UD-Q6_K_XL 237.29 0.99926049 0.99767995 0.99853696 96.60% 96.40% 100% / 100% 0.99417562 0.99165804

Q5 and Q6 are quality upgrades over our Q4 HQ. Comparisons with Unsloth are metric-specific: no universal or all-metric superiority claim is made. Q5's combined fidelity is slightly below Unsloth Q5 while its text Top-5 is higher and code Top-5 lower; the measured table is authoritative.

Verified Size Comparison

Tier Sakura HQ MiB Unsloth same tier MiB Sakura size advantage
Q4 169.79 167.54 2.25 MiB / 1.34% larger
Q5 199.88 200.32 0.45 MiB / 0.22% smaller
Q6 231.75 237.29 5.53 MiB / 2.33% smaller

Q4 HQ remains 28.45% smaller than Unsloth UD-Q6_K_XL and 15.24% smaller than UD-Q5_K_XL; these cross-tier comparisons are size statements, not equivalent-quality claims. Only standalone text GGUF files are compared, excluding multimodal mmproj files. Official Unsloth files.

Native AutoRound Configuration & Provenance

All variants use AutoRound 0.16.0 native SignRoundV2 (enable_alg_ext=True), 50 iterations, 256 independent synthetic retrieval samples, seqlen 128, batch size 1, seed 42, CPU. Samples: 48 EN queries, 48 EN docs, 48 DE queries, 48 DE docs, 32 code queries and 32 code snippets; zero exact benchmark overlap.

Tier Optimized transformer tensors Uncalibrated token tensor Central PLE Output head Norms/scalars Total tensor types
Q4 HQ 168 Q4_K + 48 block-PLE Q8_0 Q4_K Q8_0 Q6_K 194 F32 169 Q4_K, 49 Q8_0, 1 Q6_K, 194 F32
Q5 HQ 168 Q5_K + 48 block-PLE Q8_0 Q5_K Q8_0 Q8_0 194 F32 169 Q5_K, 50 Q8_0, 194 F32
Q6 HQ 168 Q6_K + 48 block-PLE Q8_0 Q6_K Q8_0 Q8_0 194 F32 169 Q6_K, 50 Q8_0, 194 F32

Each variant optimizes in its matching native scheme, then directly packs its optimized scale/minimum/double-scale state. All 216 transformer linears pass source-type, attribute and numerical roundtrip checks. No llama-quantize step and no GGUF dequantize/requantize training input are used. Embedding-path tensors outside the 216 linears are packed directly from the pinned BF16 source. Q5 uses a project-local correction to AutoRound's Q5_K packer: mask both low nibbles before combining 5-bit codes, and use bits=5 in the uncalibrated token path. Seven real/corner-case roundtrip tests and every full-model optimized-tensor check pass. Installed/shared AutoRound files and llama.cpp remain unchanged; the correction and source hashes are documented in the Q5 build information.

MRL Fidelity — all measured dimensions

Tier Dimension Mean Cosine Min Cosine Spearman Text Top-5 Code Top-5 Text/Code Recall@5
Q4 HQ 768 0.99400270 0.98141599 0.98857239 90.00% 87.60% 100% / 100%
Q4 HQ 512 0.99415833 0.98182827 0.98863087 90.40% 89.20% 100% / 100%
Q4 HQ 256 0.99473745 0.98378700 0.98938747 89.80% 88.00% 100% / 100%
Q4 HQ 128 0.99628752 0.98847973 0.98914106 87.40% 84.80% 100% / 100%
Q5 HQ 768 0.99828082 0.99384677 0.99664205 95.40% 92.80% 100% / 100%
Q5 HQ 512 0.99831867 0.99375683 0.99660520 94.40% 92.40% 100% / 100%
Q5 HQ 256 0.99848163 0.99414456 0.99667793 94.40% 94.40% 100% / 100%
Q5 HQ 128 0.99893159 0.99585915 0.99643579 91.40% 92.80% 100% / 100%
Q6 HQ 768 0.99940187 0.99797606 0.99879852 96.80% 97.20% 100% / 100%
Q6 HQ 512 0.99941486 0.99799049 0.99876861 96.00% 96.80% 100% / 100%
Q6 HQ 256 0.99947387 0.99815649 0.99874843 96.00% 98.00% 100% / 100%
Q6 HQ 128 0.99962199 0.99861580 0.99852872 96.80% 95.60% 100% / 100%

Usage Instructions

Requirements

You need a recent build of llama.cpp that includes PR #30054 (merged October 2026, commit 4fbc76d or newer). Older builds will return unknown model architecture: 'gemma-embedding2'.

1. Running via llama-server (Recommended)

Start the embedding server:


llama-server --embedding --pooling mean -m sakura-embeddinggemma-2-autoround-v2-Q5_K_S-HQ-cal256.gguf --port 8080

Query via standard OpenAI-compatible API or native HTTP endpoint:


import requests

import numpy as np

# Native llama.cpp endpoint

response = requests.post(

    "http://127.0.0.1:8080/embedding",

    json={"content": "Sakura quantized embedding models provide high retrieval accuracy."}

)

embedding = response.json()[0]["embedding"]

while isinstance(embedding, list) and isinstance(embedding[0], list):

    embedding = embedding[0]

vec = np.array(embedding, dtype=np.float32)

# L2-normalize

vec = vec / np.linalg.norm(vec)

print(f"Embedding dimension: {vec.shape[0]}, norm: {np.linalg.norm(vec):.4f}")

2. Running via llama-cli


llama-cli --embedding -m sakura-embeddinggemma-2-autoround-v2-Q5_K_S-HQ-cal256.gguf -p "A query to embed"

Runtime verified with upstream llama.cpp commit 51ce9c11a6f2dfa895696c0048c4333e8953b728, including PR #30054. Smoke tests cover server loading, finite 768d embeddings for English, German and code, and all four MRL dimensions. No llama.cpp patches are required.

Reproducibility & Limits

Full quality benchmark JSON includes all prescribed metrics and EN/DE/code views. Q5 build, Q6 build, Q5 tensor map, Q6 tensor map, SHA256 checksums. Original Q4 evidence: build, benchmark, tensor map. This is a small BF16-fidelity benchmark, not proof of general real-world retrieval superiority. nDCG grades agreement with the BF16 reference ranking, not independent human relevance judgments. Runtime smoke verifies server load, finite 768d EN/DE/code embeddings and MRL 768/512/256/128 with clean llama.cpp commit 51ce9c11a6f2dfa895696c0048c4333e8953b728 (PR #30054).

License & Attribution

Base model: Gemma Terms of Use, Google. Quantization: Intel AutoRound 0.16.0 / native SignRoundV2. Format/runtime: GGML / llama.cpp. Publisher: Sakura Quant Team (webmp3).


中文说明 · 樱花 (Simplified Chinese)

English above. 本节为上文的中文翻译(Sakura = 樱花 yīnghuā);完整的独立中文版见 README_zh.md。

Sakura EmbeddingGemma 2 — AutoRound V2 HQ GGUF

⭐ 选择你的 HQ 质量档位

这三个都是自定义的原生 AutoRound 混合精度文本嵌入模型,不是完整的 Q8 模型,也不是未经修改的 S/M 预设。文件名标识的是基础存储族;选定的 PLE 张量使用 Q8_0。Q5/Q6 还使用 Q8_0 的输出嵌入头。

基础模型:google/embeddinggemma-2,固定快照 914f7f89142e33e77833254d9c9b90c3cef7303b。架构 gemma-embedding2,413 个张量,文本路径约 271M 参数。原生嵌入 768 维;截断并重新做 L2 归一化后支持 MRL 512/256/128。

文件详情

档位 文件 字节数 MiB SHA256
推荐的平衡之选:Q5 HQ sakura-embeddinggemma-2-autoround-v2-Q5_K_S-HQ-cal256.gguf 209,585,440 199.88 d0b5f556ed4916425a3512e4af1f0d7e631f9f85b6224b6b6d956c8f33adfaa0
实测保真度最高:Q6 HQ sakura-embeddinggemma-2-autoround-v2-Q6_K-HQ-cal256.gguf 243,008,800 231.75 e5c8680e1562e47ac5d379a25dc98972ccbd037df858dbca387c1df316ae0430
紧凑:Q4 HQ sakura-embeddinggemma-2-autoround-v2-Q4_K_S-HQ-cal256.gguf 178,033,120 169.79 e01262207bf83d43043f9c880d4388b8b35f002c3fdaf86101252b3356f62ab1

并排比较 — 相同的 BF16 基准,768 维

Mean Cosine 和 Spearman 合并了 300 个向量:100 个文本查询、100 个文本文档、50 个代码查询和 50 个代码片段。文本和代码的排序指标分别列出。所有行都使用相同的现有参照、前缀、指标函数和运行时参数。

变体 MiB Mean Cosine Min Cosine Spearman Text Top-5 Code Top-5 Text/Code Recall@5 Text nDCG@10 Code nDCG@10
Sakura Q4 HQ 169.79 0.99400270 0.98141599 0.98857239 90.00% 87.60% 100% / 100% 0.97668917 0.97540851
Unsloth UD-Q4_K_XL 167.54 0.99354154 0.97681332 0.98770997 89.60% 88.40% 100% / 100% 0.97352332 0.97615687
Sakura Q5 HQ 199.88 0.99828082 0.99384677 0.99664205 95.40% 92.80% 100% / 100% 0.98892991 0.99019274
Unsloth UD-Q5_K_XL 200.32 0.99833661 0.99317700 0.99675520 94.80% 94.40% 100% / 100% 0.98853356 0.99009068
Sakura Q6 HQ 231.75 0.99940187 0.99797606 0.99879852 96.80% 97.20% 100% / 100% 0.99555133 0.99365559
Unsloth UD-Q6_K_XL 237.29 0.99926049 0.99767995 0.99853696 96.60% 96.40% 100% / 100% 0.99417562 0.99165804

Q5 和 Q6 是相对于我们的 Q4 HQ 的质量升级。与 Unsloth 的比较是针对具体指标的:我们不作任何普遍的或所有指标上都更优的声明。Q5 的综合保真度略低于 Unsloth Q5,同时它的 text Top-5 更高、code Top-5 更低;以实测表格为准。

已核实的大小比较

档位 Sakura HQ MiB Unsloth 同档 MiB Sakura 的大小优势
Q4 169.79 167.54 大 2.25 MiB / 1.34%
Q5 199.88 200.32 小 0.45 MiB / 0.22%
Q6 231.75 237.29 小 5.53 MiB / 2.33%

Q4 HQ 仍然比 Unsloth UD-Q6_K_XL 小 28.45%,比 UD-Q5_K_XL 小 15.24%;这些跨档位的比较是大小方面的陈述,不是同等质量的声明。只比较独立的文本 GGUF 文件,不包括多模态 mmproj 文件。Unsloth 官方文件。

原生 AutoRound 配置与来源

所有变体都使用 AutoRound 0.16.0 原生 SignRoundV2(enable_alg_ext=True)、50 次迭代、256 个独立的合成检索样本、序列长度 128、批大小 1、随机种子 42、CPU。样本:48 个英文查询、48 个英文文档、48 个德文查询、48 个德文文档、32 个代码查询和 32 个代码片段;与基准完全没有重合。

档位 经优化的 Transformer 张量 未校准的 token 张量 中央 PLE 输出头 归一化/标量 张量类型合计
Q4 HQ 168 Q4_K + 48 block-PLE Q8_0 Q4_K Q8_0 Q6_K 194 F32 169 Q4_K, 49 Q8_0, 1 Q6_K, 194 F32
Q5 HQ 168 Q5_K + 48 block-PLE Q8_0 Q5_K Q8_0 Q8_0 194 F32 169 Q5_K, 50 Q8_0, 194 F32
Q6 HQ 168 Q6_K + 48 block-PLE Q8_0 Q6_K Q8_0 Q8_0 194 F32 169 Q6_K, 50 Q8_0, 194 F32

每个变体都在与其匹配的原生方案下优化,然后直接打包其优化后的 scale/minimum/double-scale 状态。全部 216 个 Transformer 线性层都通过了源类型、属性和数值往返检查。没有使用 llama-quantize 步骤,也没有使用 GGUF 反量化/重新量化的训练输入。216 个线性层之外的嵌入路径张量直接从固定的 BF16 源文件打包。 Q5 使用了对 AutoRound 的 Q5_K 打包器的项目本地修正:在合并 5 比特编码之前屏蔽两个低半字节,并在未校准的 token 路径中使用 bits=5。七个真实/边界情况的往返测试以及每一项完整模型的优化张量检查都通过。已安装/共享的 AutoRound 文件和 llama.cpp 保持不变;该修正和源文件哈希记录在 Q5 构建信息中。

MRL 保真度 — 所有测量过的维度

档位 维度 Mean Cosine Min Cosine Spearman Text Top-5 Code Top-5 Text/Code Recall@5
Q4 HQ 768 0.99400270 0.98141599 0.98857239 90.00% 87.60% 100% / 100%
Q4 HQ 512 0.99415833 0.98182827 0.98863087 90.40% 89.20% 100% / 100%
Q4 HQ 256 0.99473745 0.98378700 0.98938747 89.80% 88.00% 100% / 100%
Q4 HQ 128 0.99628752 0.98847973 0.98914106 87.40% 84.80% 100% / 100%
Q5 HQ 768 0.99828082 0.99384677 0.99664205 95.40% 92.80% 100% / 100%
Q5 HQ 512 0.99831867 0.99375683 0.99660520 94.40% 92.40% 100% / 100%
Q5 HQ 256 0.99848163 0.99414456 0.99667793 94.40% 94.40% 100% / 100%
Q5 HQ 128 0.99893159 0.99585915 0.99643579 91.40% 92.80% 100% / 100%
Q6 HQ 768 0.99940187 0.99797606 0.99879852 96.80% 97.20% 100% / 100%
Q6 HQ 512 0.99941486 0.99799049 0.99876861 96.00% 96.80% 100% / 100%
Q6 HQ 256 0.99947387 0.99815649 0.99874843 96.00% 98.00% 100% / 100%
Q6 HQ 128 0.99962199 0.99861580 0.99852872 96.80% 95.60% 100% / 100%

使用说明

要求

你需要一个包含 PR #30054(2026 年 10 月合并,提交 4fbc76d 或更新)的较新 llama.cpp 构建。更旧的构建会返回 unknown model architecture: 'gemma-embedding2'。

1. 通过 llama-server 运行(推荐)

启动嵌入服务器:


llama-server --embedding --pooling mean -m sakura-embeddinggemma-2-autoround-v2-Q5_K_S-HQ-cal256.gguf --port 8080

通过标准的 OpenAI 兼容 API 或原生 HTTP 端点查询:


import requests

import numpy as np

# Native llama.cpp endpoint

response = requests.post(

    "http://127.0.0.1:8080/embedding",

    json={"content": "Sakura quantized embedding models provide high retrieval accuracy."}

)

embedding = response.json()[0]["embedding"]

while isinstance(embedding, list) and isinstance(embedding[0], list):

    embedding = embedding[0]

vec = np.array(embedding, dtype=np.float32)

# L2-normalize

vec = vec / np.linalg.norm(vec)

print(f"Embedding dimension: {vec.shape[0]}, norm: {np.linalg.norm(vec):.4f}")

2. 通过 llama-cli 运行


llama-cli --embedding -m sakura-embeddinggemma-2-autoround-v2-Q5_K_S-HQ-cal256.gguf -p "A query to embed"

运行时已用上游 llama.cpp 提交 51ce9c11a6f2dfa895696c0048c4333e8953b728(包含 PR #30054)验证。冒烟测试涵盖服务器加载、英文、德文和代码的有限 768 维嵌入,以及全部四种 MRL 维度。不需要任何 llama.cpp 补丁。

可复现性与局限

完整的质量基准 JSON 包含所有规定的指标以及 EN/DE/code 视图。Q5 构建、Q6 构建、Q5 张量映射、Q6 张量映射、SHA256 校验和。原始 Q4 证据:构建、基准、张量映射。 这是一个小型的 BF16 保真度基准,不能证明在真实世界的检索中整体更优。nDCG 衡量的是与 BF16 参照排序的一致程度,而不是独立的人工相关性判断。运行时冒烟测试验证了服务器加载、有限的 768 维 EN/DE/code 嵌入以及 MRL 768/512/256/128,使用的是干净的 llama.cpp 提交 51ce9c11a6f2dfa895696c0048c4333e8953b728(PR #30054)。

许可证与归属

基础模型:Google 的 Gemma 使用条款。量化:Intel AutoRound 0.16.0 / 原生 SignRoundV2。格式/运行时:GGML / llama.cpp。发布者:Sakura Quant Team(webmp3)。

Downloads last month
510
GGUF
Model size
0.3B params
Architecture
gemma-embedding2
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF

Quantized
(62)
this model