parmanu-lcs2/gemma-4-9B

This model is a pruned version of gemma-4-9B, compressed using TRIDENT (paper: coming soon).

Pruning details

Setting Value
Pruning method TRIDENT (paper: coming soon)
Target compression ratio 25%
Parameters after pruning 8,969,794,864
Training data SlimOrca (10,000 samples)

Hardware

All pruning was performed on an NVIDIA A100 GPU.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("parmanu-lcs2/gemma-4-12B", trust_remote_code=True, dtype="auto")
tokenizer = AutoTokenizer.from_pretrained("parmanu-lcs2/gemma-4-12B")

trust_remote_code=True is required because pruning leaves each layer's MLP with a different width. This architecture is defined in the bundled modeling_trident.py.

References

  • Original model: gemma-4-12B
  • TRIDENT paper: Coming soon
Downloads last month
-
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train parmanu-lcs2/gemma-4-9B