Instructions to use webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF:Q4_K_S # Run inference directly in the terminal: llama cli -hf webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF:Q4_K_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF:Q4_K_S # Run inference directly in the terminal: llama cli -hf webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF:Q4_K_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF:Q4_K_S # Run inference directly in the terminal: ./llama-cli -hf webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF:Q4_K_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF:Q4_K_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF:Q4_K_S
Use Docker
docker model run hf.co/webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF:Q4_K_S
- LM Studio
- Jan
- Ollama
How to use webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF with Ollama:
ollama run hf.co/webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF:Q4_K_S
- Unsloth Desktop
- Docker Model Runner
How to use webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF with Docker Model Runner:
docker model run hf.co/webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF:Q4_K_S
- Lemonade
How to use webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF:Q4_K_S
Run and chat with the model
lemonade run user.Sakura-EmbeddingGemma-2-AutoRound-GGUF-Q4_K_S
List all available models
lemonade list
- Atomic Chat
- Sakura EmbeddingGemma 2 — AutoRound V2 HQ GGUF
- Sakura EmbeddingGemma 2 — AutoRound V2 HQ GGUF
Sakura EmbeddingGemma 2 — AutoRound V2 HQ GGUF
⭐ Choose your HQ quality tier
- Q5 HQ — recommended balance: substantially higher BF16 fidelity than Q4 for a modest size increase.
- Q6 HQ — highest measured fidelity: the best BF16 alignment among our three tested HQ variants.
- Q4 HQ — compact: the smallest HQ option, with its original measured results and bytes unchanged.
All three are custom native AutoRound mixed-precision text-embedding models, not full Q8 models or unmodified S/M presets. The filename identifies the base storage family; selected PLE tensors use Q8_0. Q5/Q6 also use a Q8_0 output embedding head.
Base model: google/embeddinggemma-2, pinned snapshot 914f7f89142e33e77833254d9c9b90c3cef7303b. Architecture gemma-embedding2, 413 tensors, approximately 271M text-path parameters. Native embeddings 768d; MRL 512/256/128 after truncation and L2 renormalization.
File Details
| Tier | File | Bytes | MiB | SHA256 |
|---|---|---|---|---|
| Recommended balance: Q5 HQ | sakura-embeddinggemma-2-autoround-v2-Q5_K_S-HQ-cal256.gguf | 209,585,440 | 199.88 | d0b5f556ed4916425a3512e4af1f0d7e631f9f85b6224b6b6d956c8f33adfaa0 |
| Highest measured fidelity: Q6 HQ | sakura-embeddinggemma-2-autoround-v2-Q6_K-HQ-cal256.gguf | 243,008,800 | 231.75 | e5c8680e1562e47ac5d379a25dc98972ccbd037df858dbca387c1df316ae0430 |
| Compact: Q4 HQ | sakura-embeddinggemma-2-autoround-v2-Q4_K_S-HQ-cal256.gguf | 178,033,120 | 169.79 | e01262207bf83d43043f9c880d4388b8b35f002c3fdaf86101252b3356f62ab1 |
Side-by-Side Comparison — same BF16 benchmark, 768d
Mean Cosine and Spearman combine 300 vectors: 100 text queries, 100 text documents, 50 code queries and 50 code snippets. Text and code ranking metrics are shown separately. All rows use the same existing reference, prefixes, metric functions and runtime flags.
| Variant | MiB | Mean Cosine | Min Cosine | Spearman | Text Top-5 | Code Top-5 | Text/Code Recall@5 | Text nDCG@10 | Code nDCG@10 |
|---|---|---|---|---|---|---|---|---|---|
| Sakura Q4 HQ | 169.79 | 0.99400270 | 0.98141599 | 0.98857239 | 90.00% | 87.60% | 100% / 100% | 0.97668917 | 0.97540851 |
| Unsloth UD-Q4_K_XL | 167.54 | 0.99354154 | 0.97681332 | 0.98770997 | 89.60% | 88.40% | 100% / 100% | 0.97352332 | 0.97615687 |
| Sakura Q5 HQ | 199.88 | 0.99828082 | 0.99384677 | 0.99664205 | 95.40% | 92.80% | 100% / 100% | 0.98892991 | 0.99019274 |
| Unsloth UD-Q5_K_XL | 200.32 | 0.99833661 | 0.99317700 | 0.99675520 | 94.80% | 94.40% | 100% / 100% | 0.98853356 | 0.99009068 |
| Sakura Q6 HQ | 231.75 | 0.99940187 | 0.99797606 | 0.99879852 | 96.80% | 97.20% | 100% / 100% | 0.99555133 | 0.99365559 |
| Unsloth UD-Q6_K_XL | 237.29 | 0.99926049 | 0.99767995 | 0.99853696 | 96.60% | 96.40% | 100% / 100% | 0.99417562 | 0.99165804 |
Q5 and Q6 are quality upgrades over our Q4 HQ. Comparisons with Unsloth are metric-specific: no universal or all-metric superiority claim is made. Q5's combined fidelity is slightly below Unsloth Q5 while its text Top-5 is higher and code Top-5 lower; the measured table is authoritative.
Verified Size Comparison
| Tier | Sakura HQ MiB | Unsloth same tier MiB | Sakura size advantage |
|---|---|---|---|
| Q4 | 169.79 | 167.54 | 2.25 MiB / 1.34% larger |
| Q5 | 199.88 | 200.32 | 0.45 MiB / 0.22% smaller |
| Q6 | 231.75 | 237.29 | 5.53 MiB / 2.33% smaller |
Q4 HQ remains 28.45% smaller than Unsloth UD-Q6_K_XL and 15.24% smaller than UD-Q5_K_XL; these cross-tier comparisons are size statements, not equivalent-quality claims. Only standalone text GGUF files are compared, excluding multimodal mmproj files. Official Unsloth files.
Native AutoRound Configuration & Provenance
All variants use AutoRound 0.16.0 native SignRoundV2 (enable_alg_ext=True), 50 iterations, 256 independent synthetic retrieval samples, seqlen 128, batch size 1, seed 42, CPU. Samples: 48 EN queries, 48 EN docs, 48 DE queries, 48 DE docs, 32 code queries and 32 code snippets; zero exact benchmark overlap.
| Tier | Optimized transformer tensors | Uncalibrated token tensor | Central PLE | Output head | Norms/scalars | Total tensor types |
|---|---|---|---|---|---|---|
| Q4 HQ | 168 Q4_K + 48 block-PLE Q8_0 | Q4_K | Q8_0 | Q6_K | 194 F32 | 169 Q4_K, 49 Q8_0, 1 Q6_K, 194 F32 |
| Q5 HQ | 168 Q5_K + 48 block-PLE Q8_0 | Q5_K | Q8_0 | Q8_0 | 194 F32 | 169 Q5_K, 50 Q8_0, 194 F32 |
| Q6 HQ | 168 Q6_K + 48 block-PLE Q8_0 | Q6_K | Q8_0 | Q8_0 | 194 F32 | 169 Q6_K, 50 Q8_0, 194 F32 |
Each variant optimizes in its matching native scheme, then directly packs its optimized scale/minimum/double-scale state. All 216 transformer linears pass source-type, attribute and numerical roundtrip checks. No llama-quantize step and no GGUF dequantize/requantize training input are used. Embedding-path tensors outside the 216 linears are packed directly from the pinned BF16 source. Q5 uses a project-local correction to AutoRound's Q5_K packer: mask both low nibbles before combining 5-bit codes, and use bits=5 in the uncalibrated token path. Seven real/corner-case roundtrip tests and every full-model optimized-tensor check pass. Installed/shared AutoRound files and llama.cpp remain unchanged; the correction and source hashes are documented in the Q5 build information.
MRL Fidelity — all measured dimensions
| Tier | Dimension | Mean Cosine | Min Cosine | Spearman | Text Top-5 | Code Top-5 | Text/Code Recall@5 |
|---|---|---|---|---|---|---|---|
| Q4 HQ | 768 | 0.99400270 | 0.98141599 | 0.98857239 | 90.00% | 87.60% | 100% / 100% |
| Q4 HQ | 512 | 0.99415833 | 0.98182827 | 0.98863087 | 90.40% | 89.20% | 100% / 100% |
| Q4 HQ | 256 | 0.99473745 | 0.98378700 | 0.98938747 | 89.80% | 88.00% | 100% / 100% |
| Q4 HQ | 128 | 0.99628752 | 0.98847973 | 0.98914106 | 87.40% | 84.80% | 100% / 100% |
| Q5 HQ | 768 | 0.99828082 | 0.99384677 | 0.99664205 | 95.40% | 92.80% | 100% / 100% |
| Q5 HQ | 512 | 0.99831867 | 0.99375683 | 0.99660520 | 94.40% | 92.40% | 100% / 100% |
| Q5 HQ | 256 | 0.99848163 | 0.99414456 | 0.99667793 | 94.40% | 94.40% | 100% / 100% |
| Q5 HQ | 128 | 0.99893159 | 0.99585915 | 0.99643579 | 91.40% | 92.80% | 100% / 100% |
| Q6 HQ | 768 | 0.99940187 | 0.99797606 | 0.99879852 | 96.80% | 97.20% | 100% / 100% |
| Q6 HQ | 512 | 0.99941486 | 0.99799049 | 0.99876861 | 96.00% | 96.80% | 100% / 100% |
| Q6 HQ | 256 | 0.99947387 | 0.99815649 | 0.99874843 | 96.00% | 98.00% | 100% / 100% |
| Q6 HQ | 128 | 0.99962199 | 0.99861580 | 0.99852872 | 96.80% | 95.60% | 100% / 100% |
Usage Instructions
Requirements
You need a recent build of llama.cpp that includes PR #30054 (merged October 2026, commit 4fbc76d or newer). Older builds will return unknown model architecture: 'gemma-embedding2'.
1. Running via llama-server (Recommended)
Start the embedding server:
llama-server --embedding --pooling mean -m sakura-embeddinggemma-2-autoround-v2-Q5_K_S-HQ-cal256.gguf --port 8080
Query via standard OpenAI-compatible API or native HTTP endpoint:
import requests
import numpy as np
# Native llama.cpp endpoint
response = requests.post(
"http://127.0.0.1:8080/embedding",
json={"content": "Sakura quantized embedding models provide high retrieval accuracy."}
)
embedding = response.json()[0]["embedding"]
while isinstance(embedding, list) and isinstance(embedding[0], list):
embedding = embedding[0]
vec = np.array(embedding, dtype=np.float32)
# L2-normalize
vec = vec / np.linalg.norm(vec)
print(f"Embedding dimension: {vec.shape[0]}, norm: {np.linalg.norm(vec):.4f}")
2. Running via llama-cli
llama-cli --embedding -m sakura-embeddinggemma-2-autoround-v2-Q5_K_S-HQ-cal256.gguf -p "A query to embed"
Runtime verified with upstream llama.cpp commit 51ce9c11a6f2dfa895696c0048c4333e8953b728, including PR #30054. Smoke tests cover server loading, finite 768d embeddings for English, German and code, and all four MRL dimensions. No llama.cpp patches are required.
Reproducibility & Limits
Full quality benchmark JSON includes all prescribed metrics and EN/DE/code views. Q5 build, Q6 build, Q5 tensor map, Q6 tensor map, SHA256 checksums. Original Q4 evidence: build, benchmark, tensor map.
This is a small BF16-fidelity benchmark, not proof of general real-world retrieval superiority. nDCG grades agreement with the BF16 reference ranking, not independent human relevance judgments. Runtime smoke verifies server load, finite 768d EN/DE/code embeddings and MRL 768/512/256/128 with clean llama.cpp commit 51ce9c11a6f2dfa895696c0048c4333e8953b728 (PR #30054).
License & Attribution
Base model: Gemma Terms of Use, Google. Quantization: Intel AutoRound 0.16.0 / native SignRoundV2. Format/runtime: GGML / llama.cpp. Publisher: Sakura Quant Team (webmp3).
中文说明 · 樱花 (Simplified Chinese)
English above. 本节为上文的中文翻译(Sakura = 樱花 yīnghuā);完整的独立中文版见 README_zh.md。
Sakura EmbeddingGemma 2 — AutoRound V2 HQ GGUF
⭐ 选择你的 HQ 质量档位
- Q5 HQ — 推荐的平衡之选:相比 Q4,只增加少量体积,就获得明显更高的 BF16 保真度。
- Q6 HQ — 实测保真度最高:在我们测试的三个 HQ 变体中,与 BF16 的对齐程度最好。
- Q4 HQ — 紧凑:最小的 HQ 选项,其原有的实测结果和字节保持不变。
这三个都是自定义的原生 AutoRound 混合精度文本嵌入模型,不是完整的 Q8 模型,也不是未经修改的 S/M 预设。文件名标识的是基础存储族;选定的 PLE 张量使用 Q8_0。Q5/Q6 还使用 Q8_0 的输出嵌入头。
基础模型:google/embeddinggemma-2,固定快照 914f7f89142e33e77833254d9c9b90c3cef7303b。架构 gemma-embedding2,413 个张量,文本路径约 271M 参数。原生嵌入 768 维;截断并重新做 L2 归一化后支持 MRL 512/256/128。
文件详情
| 档位 | 文件 | 字节数 | MiB | SHA256 |
|---|---|---|---|---|
| 推荐的平衡之选:Q5 HQ | sakura-embeddinggemma-2-autoround-v2-Q5_K_S-HQ-cal256.gguf | 209,585,440 | 199.88 | d0b5f556ed4916425a3512e4af1f0d7e631f9f85b6224b6b6d956c8f33adfaa0 |
| 实测保真度最高:Q6 HQ | sakura-embeddinggemma-2-autoround-v2-Q6_K-HQ-cal256.gguf | 243,008,800 | 231.75 | e5c8680e1562e47ac5d379a25dc98972ccbd037df858dbca387c1df316ae0430 |
| 紧凑:Q4 HQ | sakura-embeddinggemma-2-autoround-v2-Q4_K_S-HQ-cal256.gguf | 178,033,120 | 169.79 | e01262207bf83d43043f9c880d4388b8b35f002c3fdaf86101252b3356f62ab1 |
并排比较 — 相同的 BF16 基准,768 维
Mean Cosine 和 Spearman 合并了 300 个向量:100 个文本查询、100 个文本文档、50 个代码查询和 50 个代码片段。文本和代码的排序指标分别列出。所有行都使用相同的现有参照、前缀、指标函数和运行时参数。
| 变体 | MiB | Mean Cosine | Min Cosine | Spearman | Text Top-5 | Code Top-5 | Text/Code Recall@5 | Text nDCG@10 | Code nDCG@10 |
|---|---|---|---|---|---|---|---|---|---|
| Sakura Q4 HQ | 169.79 | 0.99400270 | 0.98141599 | 0.98857239 | 90.00% | 87.60% | 100% / 100% | 0.97668917 | 0.97540851 |
| Unsloth UD-Q4_K_XL | 167.54 | 0.99354154 | 0.97681332 | 0.98770997 | 89.60% | 88.40% | 100% / 100% | 0.97352332 | 0.97615687 |
| Sakura Q5 HQ | 199.88 | 0.99828082 | 0.99384677 | 0.99664205 | 95.40% | 92.80% | 100% / 100% | 0.98892991 | 0.99019274 |
| Unsloth UD-Q5_K_XL | 200.32 | 0.99833661 | 0.99317700 | 0.99675520 | 94.80% | 94.40% | 100% / 100% | 0.98853356 | 0.99009068 |
| Sakura Q6 HQ | 231.75 | 0.99940187 | 0.99797606 | 0.99879852 | 96.80% | 97.20% | 100% / 100% | 0.99555133 | 0.99365559 |
| Unsloth UD-Q6_K_XL | 237.29 | 0.99926049 | 0.99767995 | 0.99853696 | 96.60% | 96.40% | 100% / 100% | 0.99417562 | 0.99165804 |
Q5 和 Q6 是相对于我们的 Q4 HQ 的质量升级。与 Unsloth 的比较是针对具体指标的:我们不作任何普遍的或所有指标上都更优的声明。Q5 的综合保真度略低于 Unsloth Q5,同时它的 text Top-5 更高、code Top-5 更低;以实测表格为准。
已核实的大小比较
| 档位 | Sakura HQ MiB | Unsloth 同档 MiB | Sakura 的大小优势 |
|---|---|---|---|
| Q4 | 169.79 | 167.54 | 大 2.25 MiB / 1.34% |
| Q5 | 199.88 | 200.32 | 小 0.45 MiB / 0.22% |
| Q6 | 231.75 | 237.29 | 小 5.53 MiB / 2.33% |
Q4 HQ 仍然比 Unsloth UD-Q6_K_XL 小 28.45%,比 UD-Q5_K_XL 小 15.24%;这些跨档位的比较是大小方面的陈述,不是同等质量的声明。只比较独立的文本 GGUF 文件,不包括多模态 mmproj 文件。Unsloth 官方文件。
原生 AutoRound 配置与来源
所有变体都使用 AutoRound 0.16.0 原生 SignRoundV2(enable_alg_ext=True)、50 次迭代、256 个独立的合成检索样本、序列长度 128、批大小 1、随机种子 42、CPU。样本:48 个英文查询、48 个英文文档、48 个德文查询、48 个德文文档、32 个代码查询和 32 个代码片段;与基准完全没有重合。
| 档位 | 经优化的 Transformer 张量 | 未校准的 token 张量 | 中央 PLE | 输出头 | 归一化/标量 | 张量类型合计 |
|---|---|---|---|---|---|---|
| Q4 HQ | 168 Q4_K + 48 block-PLE Q8_0 | Q4_K | Q8_0 | Q6_K | 194 F32 | 169 Q4_K, 49 Q8_0, 1 Q6_K, 194 F32 |
| Q5 HQ | 168 Q5_K + 48 block-PLE Q8_0 | Q5_K | Q8_0 | Q8_0 | 194 F32 | 169 Q5_K, 50 Q8_0, 194 F32 |
| Q6 HQ | 168 Q6_K + 48 block-PLE Q8_0 | Q6_K | Q8_0 | Q8_0 | 194 F32 | 169 Q6_K, 50 Q8_0, 194 F32 |
每个变体都在与其匹配的原生方案下优化,然后直接打包其优化后的 scale/minimum/double-scale 状态。全部 216 个 Transformer 线性层都通过了源类型、属性和数值往返检查。没有使用 llama-quantize 步骤,也没有使用 GGUF 反量化/重新量化的训练输入。216 个线性层之外的嵌入路径张量直接从固定的 BF16 源文件打包。 Q5 使用了对 AutoRound 的 Q5_K 打包器的项目本地修正:在合并 5 比特编码之前屏蔽两个低半字节,并在未校准的 token 路径中使用 bits=5。七个真实/边界情况的往返测试以及每一项完整模型的优化张量检查都通过。已安装/共享的 AutoRound 文件和 llama.cpp 保持不变;该修正和源文件哈希记录在 Q5 构建信息中。
MRL 保真度 — 所有测量过的维度
| 档位 | 维度 | Mean Cosine | Min Cosine | Spearman | Text Top-5 | Code Top-5 | Text/Code Recall@5 |
|---|---|---|---|---|---|---|---|
| Q4 HQ | 768 | 0.99400270 | 0.98141599 | 0.98857239 | 90.00% | 87.60% | 100% / 100% |
| Q4 HQ | 512 | 0.99415833 | 0.98182827 | 0.98863087 | 90.40% | 89.20% | 100% / 100% |
| Q4 HQ | 256 | 0.99473745 | 0.98378700 | 0.98938747 | 89.80% | 88.00% | 100% / 100% |
| Q4 HQ | 128 | 0.99628752 | 0.98847973 | 0.98914106 | 87.40% | 84.80% | 100% / 100% |
| Q5 HQ | 768 | 0.99828082 | 0.99384677 | 0.99664205 | 95.40% | 92.80% | 100% / 100% |
| Q5 HQ | 512 | 0.99831867 | 0.99375683 | 0.99660520 | 94.40% | 92.40% | 100% / 100% |
| Q5 HQ | 256 | 0.99848163 | 0.99414456 | 0.99667793 | 94.40% | 94.40% | 100% / 100% |
| Q5 HQ | 128 | 0.99893159 | 0.99585915 | 0.99643579 | 91.40% | 92.80% | 100% / 100% |
| Q6 HQ | 768 | 0.99940187 | 0.99797606 | 0.99879852 | 96.80% | 97.20% | 100% / 100% |
| Q6 HQ | 512 | 0.99941486 | 0.99799049 | 0.99876861 | 96.00% | 96.80% | 100% / 100% |
| Q6 HQ | 256 | 0.99947387 | 0.99815649 | 0.99874843 | 96.00% | 98.00% | 100% / 100% |
| Q6 HQ | 128 | 0.99962199 | 0.99861580 | 0.99852872 | 96.80% | 95.60% | 100% / 100% |
使用说明
要求
你需要一个包含 PR #30054(2026 年 10 月合并,提交 4fbc76d 或更新)的较新 llama.cpp 构建。更旧的构建会返回 unknown model architecture: 'gemma-embedding2'。
1. 通过 llama-server 运行(推荐)
启动嵌入服务器:
llama-server --embedding --pooling mean -m sakura-embeddinggemma-2-autoround-v2-Q5_K_S-HQ-cal256.gguf --port 8080
通过标准的 OpenAI 兼容 API 或原生 HTTP 端点查询:
import requests
import numpy as np
# Native llama.cpp endpoint
response = requests.post(
"http://127.0.0.1:8080/embedding",
json={"content": "Sakura quantized embedding models provide high retrieval accuracy."}
)
embedding = response.json()[0]["embedding"]
while isinstance(embedding, list) and isinstance(embedding[0], list):
embedding = embedding[0]
vec = np.array(embedding, dtype=np.float32)
# L2-normalize
vec = vec / np.linalg.norm(vec)
print(f"Embedding dimension: {vec.shape[0]}, norm: {np.linalg.norm(vec):.4f}")
2. 通过 llama-cli 运行
llama-cli --embedding -m sakura-embeddinggemma-2-autoround-v2-Q5_K_S-HQ-cal256.gguf -p "A query to embed"
运行时已用上游 llama.cpp 提交 51ce9c11a6f2dfa895696c0048c4333e8953b728(包含 PR #30054)验证。冒烟测试涵盖服务器加载、英文、德文和代码的有限 768 维嵌入,以及全部四种 MRL 维度。不需要任何 llama.cpp 补丁。
可复现性与局限
完整的质量基准 JSON 包含所有规定的指标以及 EN/DE/code 视图。Q5 构建、Q6 构建、Q5 张量映射、Q6 张量映射、SHA256 校验和。原始 Q4 证据:构建、基准、张量映射。
这是一个小型的 BF16 保真度基准,不能证明在真实世界的检索中整体更优。nDCG 衡量的是与 BF16 参照排序的一致程度,而不是独立的人工相关性判断。运行时冒烟测试验证了服务器加载、有限的 768 维 EN/DE/code 嵌入以及 MRL 768/512/256/128,使用的是干净的 llama.cpp 提交 51ce9c11a6f2dfa895696c0048c4333e8953b728(PR #30054)。
许可证与归属
基础模型:Google 的 Gemma 使用条款。量化:Intel AutoRound 0.16.0 / 原生 SignRoundV2。格式/运行时:GGML / llama.cpp。发布者:Sakura Quant Team(webmp3)。
- Downloads last month
- 510
4-bit
5-bit
6-bit
Model tree for webmp3/Sakura-EmbeddingGemma-2-AutoRound-GGUF
Base model
google/embeddinggemma-2