Instructions to use Aratako/Qwen1.5-MoE-2x72B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Aratako/Qwen1.5-MoE-2x72B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Aratako/Qwen1.5-MoE-2x72B", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Aratako/Qwen1.5-MoE-2x72B", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("Aratako/Qwen1.5-MoE-2x72B", trust_remote_code=True, device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Aratako/Qwen1.5-MoE-2x72B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Aratako/Qwen1.5-MoE-2x72B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Aratako/Qwen1.5-MoE-2x72B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Aratako/Qwen1.5-MoE-2x72B
- SGLang
How to use Aratako/Qwen1.5-MoE-2x72B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Aratako/Qwen1.5-MoE-2x72B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Aratako/Qwen1.5-MoE-2x72B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Aratako/Qwen1.5-MoE-2x72B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Aratako/Qwen1.5-MoE-2x72B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Aratako/Qwen1.5-MoE-2x72B with Docker Model Runner:
docker model run hf.co/Aratako/Qwen1.5-MoE-2x72B
Qwen1.5-MoE-2x72B
Description
This model is created using MoE (Mixture of Experts) through mergekit based on Qwen/Qwen1.5-72B-Chat and abacusai/Liberated-Qwen1.5-72B without further FT.
It utilizes a customized script for MoE via mergekit, which is available here.
Due to the structural modifications introduced by MoE, the use of this model requires custom modeling file and custom configuration file. When using the model, please place these files in the same folder as the model.
This model inherits the the tongyi-qianwen license.
Benchmark
The benchmark score of the mt-bench for this model and the two base models are as follows:
1-turn, 4-bit quantization
| Model | Size | Coding | Extraction | Humanities | Math | Reasoning | Roleplay | STEM | Writing | avg_score |
|---|---|---|---|---|---|---|---|---|---|---|
| Liberated-Qwen1.5-72B | 72B | 5.8 | 7.9 | 9.6 | 6.7 | 7.0 | 9.05 | 9.55 | 9.9 | 8.1875 |
| Qwen1.5-72B-Chat | 72B | 5.5 | 8.7 | 9.7 | 8.4 | 7.5 | 9.0 | 9.45 | 9.75 | 8.5000 |
| This model | 2x72B | 5.6 | 7.8 | 9.75 | 7.0 | 8.1 | 9.0 | 9.65 | 9.8 | 8.3375 |
2-turn, 4-bit quantization
| Model | Size | Coding | Extraction | Humanities | Math | Reasoning | Roleplay | STEM | Writing | avg_score |
|---|---|---|---|---|---|---|---|---|---|---|
| Liberated-Qwen1.5-72B | 72B | 3.9 | 8.2 | 10.0 | 5.7 | 5.5 | 8.4 | 8.7 | 8.6 | 7.3750 |
| Qwen1.5-72B-Chat | 72B | 5.2 | 8.8 | 10.0 | 6.1 | 6.7 | 9.0 | 9.8 | 9.5 | 8.1375 |
| This model | 2x72B | 5.0 | 9.5 | 9.9 | 5.6 | 8.1 | 9.3 | 9.6 | 9.2 | 8.2750 |
Merge config
base_model: ./Qwen1.5-72B-Chat
gate_mode: random
dtype: bfloat16
experts:
- source_model: ./Qwen1.5-72B-Chat
positive_prompts: []
- source_model: ./Liberated-Qwen1.5-72B
positive_prompts: []
tokenizer_source: model:./Qwen1.5-72B-Chat
Gratitude
- Huge thanks to Alibaba Cloud Qwen for training and publishing the weights of Qwen model
- Thank you to abacusai for publishing fine-tuned model from Qwen
- And huge thanks to mlabonne, as I customized modeling file using phixtral as a reference
- Downloads last month
- 19

