Spark-X2.5-1.7B

SparkLLM (XHToken)

Parameters

2.0B

Architecture

Hybrid attention transformer: 1 full-attention layer combined with 3 sliding-window attention (SWA) layers, natively supporting up to 1M-token context

Released

01.09.2026

License

Apache License 2.0

Open Weights Commercial Use Multimodal BF16 Spark-X2.5 en zh multilingual (200+ languages)

Input Modalities

text

Output Modalities

text

Context (native)

1,048,576 tokens

Context (extended)

1,048,576 tokens

Openness Index Score 100.0/100

About

Spark-X2.5-1.7B is iFLYTEK's compact, general-purpose on-device language model (released by the XHToken team), designed to make capable AI practical, efficient and accessible across conversation, writing, translation, reasoning, coding, tool use and agentic workflows. It achieves leading results among open-source models of comparable size and was trained on Huawei Ascend clusters.

Its efficiency-oriented hybrid attention architecture combines one full-attention layer with three Sliding-Window Attention (SWA) layers (21 sliding + 7 full of 28 layers, sliding window 512), with Gated Attention output gates (headwise sigmoid gating, 8 Q / 2 KV heads, head dim 256, RoPE theta 5M with partial rotary 0.25 on full layers). This hybrid SWA-full attention (3:1) design substantially reduces the computational overhead of long-context models while natively supporting a context window of up to 1M tokens (1,048,576 positions). The model has ~1.7B parameters (hidden 2048, FFN 6656, vocab 131,072, tied embeddings) and supports more than 200 languages.

Pre-training used ~20 trillion tokens of web pages, books, academic publications, code and encyclopedic materials, with a dedicated long-context stage of hundreds of billions of tokens at sequence lengths up to 1M. Post-training: supervised fine-tuning, then large-scale reinforcement learning across capability domains (language understanding, reasoning, programming, tool-augmented agentic behavior, instruction following), and consolidation of the resulting domain-specialized teacher policies into one deployable model via Multi-Teacher On-Policy Distillation (MOPD).

It integrates with popular agent harnesses (Codex, Claude Code, OpenClaw, Hermes), runs on NVIDIA, Huawei, Hygon and HOUMO.AI hardware via vLLM, SGLang, llama.cpp, MLX, Ollama and LM Studio, and can be fine-tuned with LLaMA-Factory. Released September 1, 2026 under Apache 2.0.

Training Data Pretrained on ~20 trillion tokens (web pages, books, academic publications, code, encyclopedic materials) with dedicated long-context training up to 1M sequence length; post-training: supervised fine-tuning, large-scale reinforcement learning across capability domains, and MOPD consolidation of domain-specialized teacher policies. Trained on Huawei Ascend clusters. Supports 200+ languages.

Benchmark Scores

Benchmark Score Date
BFCL-V4
general_agent
39.76%
—
TAU2-Bench
general_agent
60.74%
—
TAU3-Bench
general_agent
27.92%
—
MCP-Atlas
general_agent
14.77%
—
Workspace Bench
47.66%
—
VITA-Bench
general_agent
10.10%
—
BrowseComp
general_agent
30.19%
—
SWE-bench Pro
coding_agent
13.00%
—
SWE-bench Verified
coding_agent
32.00%
—
SWE-bench Multilingual
coding_agent
21.66%
—
SciCode
36.09%
—
Gaokao 2026
61.45%
—
AIME 26
stem_reasoning
61.99%
—
HMMT Feb 26
stem_reasoning
36.14%
—
IMOAnswerBench
stem_reasoning
32.39%
—
IFEval
instruction_following
90.86%
—
IFBench
instruction_following
72.55%
—
AA-LCR
long_context
30.38%
—
Humanity's Last Exam
stem_reasoning
7.95%
—
GPQA
reasoning
46.30%
—
MCPMark
general_agent
2.30
—

Model Tree and Collection

Model tree for XHToken/Spark-X2.5-1.7B

Base model

XHToken/Spark-X2.5-1.7B-Base

Finetuned

(1)

this model

Finetunes

1 model

Quantizations

9 models

Spaces using XHToken/Spark-X2.5-1.7B 2

Collection including XHToken/Spark-X2.5-1.7B

[

Spark-X2.5

Collection

Spark-X2.5 is a compact, general-purpose language model for conversation, writing, translation, reasoning, coding, tool use, and agentic workflows. • 10 items • Updated 5 days ago • 42

](https://huggingface.co/collections/XHToken/spark-x25)

Citation

Citation

If you find our work helpful, feel free to give us a cite.

@misc{sparkx2.5,
    title  = {Spark-X2.5 4B&1.7B: Pushing the Limits of Agentic Capabilities in On-Device Models},
    author = {SparkLLM Team},
    year   = {2026}
}

Safetensors

Model size

2B params

Tensor type

BF16

·

License (Apache 2.0)

License

The Spark-X2.5 model series is licensed under the Apache 2.0 License.

Fine-Tuning (LLaMA-Factory)

Fine-Tuning

We recommend using Llama-Factory to fine-tune the model.

Quickstart (SGLang, vLLM, MLX, Ollama, LM Studio)

Quickstart

The examples below serve a local Spark-X2.5-1.7B checkpoint. Set MODEL_PATH to its absolute path before starting a container:

export MODEL_PATH=/absolute/path/to/Spark-X2.5-1.7B

SGLang

Install SGLang

Use the pre-built image that tracks the Spark-X2.5 runtime:

For NVIDIA GPUs:

docker pull lmsysorg/sglang:nightly-dev-cu13-20260827-20621aa1

For Ascend NPUs:

## A3 daily build
export SGLANG_IMAGE=quay.io/ascend/sglang:main-cann9.0.0-a3

## A2 daily build (use this instead on A2 hardware)
export SGLANG_IMAGE=quay.io/ascend/sglang:main-cann9.0.0-910b

docker pull "$SGLANG_IMAGE"

Run Inference

The following commands start an OpenAI-compatible API server configured for a maximum context length of 1,048,576 tokens. This setting requires sufficient device memory; reduce --context-length when necessary.

Server

NVIDIA GPU:

docker run --rm -it \
  --gpus '"device=0"' \
  --ipc=host \
  -p 30000:30000 \
  -v "$MODEL_PATH:/root/Spark-X2.5-1.7B:ro" \
  lmsysorg/sglang:nightly-dev-cu13-20260827-20621aa1 \
  python -m sglang.launch_server \
    --model-path /root/Spark-X2.5-1.7B \
    --served-model-name spark2.5 \
    --tool-call-parser spark25 \
    --reasoning-parser qwen3 \
    --tp-size 1 \
    --mem-fraction-static 0.8 \
    --context-length 1048576 \
    --chat-template /root/Spark-X2.5-1.7B/chat_template.jinja \
    --host 0.0.0.0 \
    --port 30000

Ascend NPU:

docker run -it --rm -e ASCEND_USE_FIA=1 --network=host --ipc=host --shm-size=16g \
    --device=/dev/davinci0 --device=/dev/davinci1 --device=/dev/davinci2 --device=/dev/davinci3 \
    --device=/dev/davinci4 --device=/dev/davinci5 --device=/dev/davinci6 --device=/dev/davinci7 \
    --device=/dev/davinci8 --device=/dev/davinci9 --device=/dev/davinci10 --device=/dev/davinci11 \
    --device=/dev/davinci12 --device=/dev/davinci13 --device=/dev/davinci14 --device=/dev/davinci15 \
    --device=/dev/davinci_manager \
    --device=/dev/devmm_svm \
    --device=/dev/hisi_hdc \
    --volume /usr/local/sbin:/usr/local/sbin \
    --volume /usr/local/Ascend/driver:/usr/local/Ascend/driver \
    --volume /usr/local/Ascend/firmware:/usr/local/Ascend/firmware \
    --volume /etc/ascend_install.info:/etc/ascend_install.info \
    --volume /var/queue_schedule:/var/queue_schedule \
    --volume ~/.cache/:/root/.cache/ \
    --volume "$MODEL_PATH:/root/Spark-X2.5-1.7B:ro" \
    --entrypoint=python \
    "$SGLANG_IMAGE" \
    -m sglang.launch_server \
      --model-path /root/Spark-X2.5-1.7B \
      --served-model-name spark2.5 \
      --tool-call-parser spark25 \
      --reasoning-parser qwen3 \
      --tp-size 1 \
      --mem-fraction-static 0.8 \
      --context-length 1048576 \
      --chat-template /root/Spark-X2.5-1.7B/chat_template.jinja \
      --host 0.0.0.0 \
      --port 30000

Client

Thinking is enabled by default by both the chat template and the Qwen3 reasoning parser. To disable thinking for a specific request, set "chat_template_kwargs": {"enable_thinking": false}.

curl -s http://localhost:30000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spark2.5",
    "messages": [
      {
        "role": "user",
        "content": "What is the capital of Anhui Province?"
      }
    ],
    "max_tokens": 131072,
    "temperature": 1,
    "top_k": -1,
    "top_p": 0.95,
    "repetition_penalty": 1,
    "presence_penalty": 0,
    "frequency_penalty": 0
  }'

vLLM

Deploy vLLM

vLLM provides an official Docker image for NVIDIA GPU deployment:

docker run --rm --gpus all \
  --ipc=host \
  -p 30000:30000 \
  -v "$MODEL_PATH:/models/Spark-X2.5-1.7B:ro" \
  vllm/vllm-openai:latest \
  --model /models/Spark-X2.5-1.7B \
  --port 30000 \
  --trust-remote-code \
  --served-model-name spark25 \
  --tensor-parallel-size 1 \
  --gpu-memory-utilization 0.7 \
  --enable-prefix-caching \
  --chat-template /models/Spark-X2.5-1.7B/chat_template.jinja

For Ascend NPUs, choose an official image for the fastest setup.

Ascend A2:

export IMAGE=quay.io/ascend/vllm-ascend:nightly-main
docker pull "$IMAGE"

export DEVICE=/dev/davinci0
export MODEL_CACHE="${HOME}/.cache"

mkdir -p "$MODEL_CACHE"

docker run --rm \
    --name vllm-ascend \
    --shm-size=1g \
    --device "$DEVICE" \
    --device /dev/davinci_manager \
    --device /dev/devmm_svm \
    --device /dev/hisi_hdc \
    -v /usr/local/dcmi:/usr/local/dcmi \
    -v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
    -v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
    -v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
    -v /etc/ascend_install.info:/etc/ascend_install.info \
    -v "$MODEL_CACHE:/root/.cache" \
    -p 8000:8000 \
    -it "$IMAGE" bash

Ascend A3:

export IMAGE=quay.io/ascend/vllm-ascend:nightly-main-a3
docker pull "$IMAGE"

export DEVICE0=/dev/davinci0
export DEVICE1=/dev/davinci1
export MODEL_CACHE="${HOME}/.cache"

mkdir -p "$MODEL_CACHE"

docker run --rm \
    --name vllm-ascend \
    --shm-size=1g \
    --device "$DEVICE0" \
    --device "$DEVICE1" \
    --device /dev/davinci_manager \
    --device /dev/devmm_svm \
    --device /dev/hisi_hdc \
    -v /usr/local/dcmi:/usr/local/dcmi \
    -v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
    -v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
    -v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
    -v /etc/ascend_install.info:/etc/ascend_install.info \
    -v "$MODEL_CACHE:/root/.cache" \
    -p 8000:8000 \
    -it "$IMAGE" bash

Ascend 950DT:

export IMAGE=quay.io/ascend/vllm-ascend:nightly-main-a5
docker pull "$IMAGE"

export MODEL_CACHE="${HOME}/.cache"

mkdir -p "$MODEL_CACHE"

docker run --rm \
    --name vllm-ascend \
    --net=host \
    --shm-size=1g \
    --device /dev/davinci0 \
    --device /dev/davinci_manager \
    --device /dev/devmm_svm \
    --device /dev/hisi_hdc \
    -v /usr/local/dcmi:/usr/local/dcmi \
    -v /usr/local/Ascend/driver/tools/hccn_tool:/usr/local/Ascend/driver/tools/hccn_tool \
    -v /usr/local/bin/npu-smi:/usr/local/bin/npu-smi \
    -v /usr/local/Ascend/driver/lib64/:/usr/local/Ascend/driver/lib64/ \
    -v /usr/local/Ascend/driver/version.info:/usr/local/Ascend/driver/version.info \
    -v /etc/ascend_install.info:/etc/ascend_install.info \
    -v "$MODEL_CACHE:/root/.cache" \
    -it "$IMAGE" bash

Install the Spark plugin inside the container:

pip install uv
uv venv ~/spark2_5
source ~/spark2_5/bin/activate
git clone https://github.com/XHToken/Spark-plugin.git
cd ./Spark-plugin
uv pip install .

Server

vllm serve "/models/Spark-X2.5-1.7B" \
 --port "30000" \
 --trust-remote-code \
 --served-model-name spark25 \
 --tensor-parallel-size 1 \
 --gpu-memory-utilization 0.7 \
 --enable-prefix-caching \
 --chat-template /models/Spark-X2.5-1.7B/chat_template.jinja

Client

curl -s http://127.0.0.1:30000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spark25",
    "messages": [{"role": "user", "content": "What is the capital of Anhui Province?"}],
    "temperature": 1.0, 
    "top_k": -1,
    "top_p": 0.95
  }'

MLX

Spark-MLX-LLM runs the original Spark-X2.5 Hugging Face checkpoints locally. It supports Apple silicon GPU, Linux CPU, and NVIDIA CUDA on Linux. No GGUF conversion is required.

Installation

git clone https://github.com/XHToken/Spark-MLX-LLM.git
cd Spark-MLX-LLM

python3 -m venv .venv
source .venv/bin/activate

## Apple silicon
python -m pip install -e .
## Linux CPU
python -m pip install -e '.[cpu]'
## Linux with CUDA 12
python -m pip install -e '.[cuda12]' 
## Linux with CUDA 13
python -m pip install -e '.[cuda13]'

Run Inference

spark-mlx-generate \
  --device gpu \
  --dtype bfloat16 \
  --model XHToken/Spark-X2.5-1.7B \
  --prompt "What is the capital of Anhui Province?" \
  --max-tokens 512 \
  --temp 0

Ollama

Build

git clone https://github.com/XHToken/llama.cpp.git llama.cpp-spark
git clone https://github.com/ollama/ollama.git ollama-spark
cd ollama-spark
export OLLAMA_LLAMA_CPP_SOURCE="$(cd ../llama.cpp-spark && pwd)"
cmake -S . -B build
cmake --build build --parallel 8

Create and Run

Create the model definition, then start the Ollama server in one terminal:

printf 'FROM /absolute/path/to/your.gguf\n' > ./Modelfile.spark
./ollama serve

Create and run the model from another terminal:

./ollama create Spark-X2.5-1.7B -f ./Modelfile.spark
./ollama run Spark-X2.5-1.7B

LM Studio

Build

git clone https://github.com/XHToken/llama.cpp.git llama.cpp-spark
cd llama.cpp-spark
cmake -S . -B build
cmake --build build --parallel 8

Set Up LM Studio

  1. Close LM Studio.

  2. Back up the selected runtime directory:

    <LM_STUDIO_HOME>/extensions/backends/<selected-runtime>/
    
    
  3. Copy the llama.cpp-spark build output into the selected runtime directory, overwriting the existing files.

  4. Place the GGUF model in the following directory:

    <LM_STUDIO_HOME>/models/<org>/<name>/
    
    

Example runtime directory on macOS:

./build/bin/* -> ~/.lmstudio/extensions/backends/llama.cpp-mac-arm64-apple-metal-advsimd-<version>/

Run with LM Studio

Open My Models, select the Spark-X2.5 model, click Load, then start a new Chat.

Run with the lms CLI

## Replace <model> with a model listed by lms ls.
lms load <model>
lms chat <model>

Benchmarks vs on-device models

Benchmarks

We evaluate our models and compare them with leading on-device models of similar size across a broad range of tasks, including agent, code, math, general and knowledge.

Benchmark Spark‑X2.5‑4B Spark‑X2.5‑1.7B Qwen3.5‑9B Qwen3.5‑4B Qwen3.5‑2B Gemma4‑12B Gemma4‑E4B Gemma4‑E2B
Agent
BFCL‑V4 65.1 46.9 66.1* 50.3* 43.6* 37.4 36.9 30.2
τ²‑bench 75.1 65.3 79.1* 79.9* 48.8* 69.0* 42.2* 24.5*
τ³‑bench 30.4 20.1 9.3 6.7 4.1 13.3 10.1 8.8
MCP‑Atlas 54.6 23.4 47.4* 40.8* 14.8 30.5* 15.0* 12.6
MCP‑Mark 14.2 2.3 13.4 12.5 – – – –
Workspace Bench 31.2 18.9 25.5 21.3 7.7 – – –
VitaBench2.0 25.2 8.3 15.6 18.2 5.2 12.4 4.8 4.4
BrowseComp 40.9 29.7 8.3 14.3 3.1 10.0 8.3 3.7
Code
SWE‑Bench Pro 44.4 10.4 33.8* 29.4* 1.9 21.9* 4.0* –
SWE‑Bench Verified 41.6 28.3 53.1* 38.8* 6.8 44.2* 14.0* –
SWE‑Bench Multilingual 53.3 23.3 43.3 27.7 5.0 32.5* – –
SciCode 34.7 18.2 32.7* 24.0 6.0 39.8 27.5 20.5
Math
Gaokao 2026 133.4 114.8 135.5 130.3 94.0 130.6 102.4 81.8
AIME 2026 90.7 69.4 88.2 83.0 30.8 82.1* 42.5* 37.5*
HMMT Feb 2026 81.2 48.4 70.8 69.7 21.5 65.6 34.2 20.5
IMO‑AnswerBench 74.2 45.4 69.8 68.5 – 57.2 26.9 22.6
General & Knowledge
IFEval 93.0 89.5 91.5* 89.8* 78.6* 94.8 45.3 34.8
IFBench 75.0 66.3 64.5 59.2 41.3* 73.5* 44.0* 22.7
AA‑LCR 56.3 24.3 63.0* 57.0* 25.6* 55.3* 34.7 18.3
HLE 12.3 6.3 14.3 8.6 2.1 13.1 3.9 2.5
GPQA 67.4 43.8 77.2 67.2 44.6 72.8 54.5 43.8
  • * denotes reported results from publicly‑released model cards / papers and - denotes scores not yet available.
  • All evaluations are conducted in thinking mode. The recommended sampling parameters for Spark-X2.5 are temperature=1.0, top_p=0.95, and top_k=-1.
  • Gaokao 2026 consists of the five 2026 Chinese GAOKAO examinations (National I,National II, Beijing, Shanghai, Tianjin), each graded out of 150 points.

Training Methods (pre-training, long-context, SFT, RL, MOPD)

Training Methods

Spark-X2.5 is pretrained on approximately 20 trillion tokens from a diverse corpus spanning web pages, books, academic publications, code, and encyclopedic materials. Particular attention is paid to data quality, domain coverage, and the sampling weights assigned to different data categories. Extensive data-mixture studies are conducted to determine an effective balance among mathematics, logic, code, and other high-value domains. This enables the models to acquire broad general knowledge while developing stronger capabilities in complex reasoning and code generation. Long-context capability is developed through a dedicated training stage comprising hundreds of billions of tokens, with sequence lengths extending to 1M tokens.

Post-training begins with supervised fine-tuning on a carefully curated corpus. This stage establishes robust instruction following, structured generation, and task-completion, while providing a stable policy initialization for reinforcement learning. We subsequently apply large-scale reinforcement learning across several capability domains, including language understanding, reasoning, programming, tool-augmented agentic behavior, and instruction following. This process yields a set of domain-specialized teacher policies, whose complementary strengths are consolidated into a single deployable model through MOPD.

Model Overview: Hybrid Attention Architecture

Model Overview

For agent tasks, balancing performance, inference speed, and cache usage has long been a key bottleneck limiting model performance. Spark-X2.5 systematically integrates and optimizes mature attention technologies, combining sliding-window attention (SWA) with a hybrid full-attention architecture. This approach leverages the strengths of both mechanisms while avoiding the limitations of relying on a single structure, achieving an effective balance among performance, inference efficiency, and KV-cache size—thereby improving its practicality and effectiveness across real-world deployment scenarios.

Spark-X2.5 hybrid architecture

Introduction and Technical Highlights

Introduction

We are introducing Spark-X2.5-4B and Spark-X2.5-1.7B, two compact, general-purpose language models designed to make capable AI more practical, efficient, and accessible. The models deliver strong performance across a broad range of everyday tasks—including conversation, writing, translation, reasoning, coding, tool use, and agentic workflows—achieving leading results among open-source models of comparable size. Spark-X2.5 combines an efficiency-oriented architecture with native context windows of up to 1M tokens, and support for more than 200 languages.

Technical Highlights:

  • Efficient Architecture and Native 1M-token Context: The models use a hybrid attention architecture that combines one full-attention layer with three sliding-window attention layers. This design substantially reduces the computational overhead typically associated with long-context models while natively supporting a context window of up to 1M tokens.
  • Strong Coding and Agent Capabilities: The models are deeply integrated with popular agent harnesses, including Codex, Claude Code, OpenClaw, and Hermes. They deliver state-of-the-art performance among models of comparable size across everyday coding, agentic workflows, reasoning, and instruction-following tasks.
  • Broad Hardware and Software Compatibility: The models support a wide range of hardware platforms, including NVIDIA, Huawei, Hygon, HOUMO.AI, etc. It is compatible with leading inference frameworks such as vLLM, SGLang, llama.cpp, MLX, and can be deployed quickly through platforms including Ollama and LM Studio. The models can also be customized using popular fine-tuning frameworks such as LLaMA-Factory. Across multiple hardware platforms, they deliver superior TTFT, TOPT, and overall inference efficiency compared with similarly sized models.
  • Advanced Training Algorithms: The models were trained on Huawei Ascend clusters. Large-scale reinforcement learning and post-training techniques such as MOPD significantly enhance its reasoning, coding, agentic, and instruction-following capabilities.

Architecture

Decoder Block ×28 input Embedding vocab 131K · d 2048 Sliding Window Attn Hybrid 8:2 · dₕ 256 · win 512 ×21 Full Attention Hybrid 8:2 · dₕ 256 · win 512 ×7 Dense FFN gelu · d 6656 Final Norm LM Head tied with embedding output
Attention
Hybrid Attention (8:2)
Layers
28
Hidden size
2048
Context
1M tokens
Parameters
2000M

Source: Hugging Face config.json · Spark2_5ForCausalLM · exact layer pattern · model repo

Training Pipeline

  1. 1
    pretraining

    Pre-training on ~20T tokens

    Diverse corpus (web pages, books, academic publications, code, encyclopedic materials) with data-mixture studies balancing mathematics, logic, code and other high-value domains.

  2. 2
    cpt

    Dedicated long-context training to 1M

    Hundreds of billions of tokens with sequence lengths extending to 1M tokens.

  3. 3
    sft

    Supervised fine-tuning

    SFT on a carefully curated corpus; establishes instruction following, structured generation and task completion; stable policy initialization for RL.

  4. 4
    rl

    Large-scale RL across capability domains

    RL across language understanding, reasoning, programming, tool-augmented agentic behavior and instruction following, yielding domain-specialized teacher policies.

  5. 5
    other

    MOPD consolidation

    Complementary strengths of the domain-specialized teacher policies consolidated into a single deployable model via Multi-Teacher On-Policy Distillation.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
Web pages pretraining — —
Books pretraining — —
Academic publications pretraining — —
Code pretraining — —
Encyclopedic materials pretraining — —

Related Models