Parameters
2.7B
Architecture
LFM2 (Hybrid)
Released
04.08.2026
License
LFM Open License v1.0
Input Modalities
Output Modalities
Context (native)
131,072 tokens
Context (extended)
131,072 tokens
About
LFM2.5-2.6B is Liquid AI's post-trained agent model - part of the LFM2.5 family of hybrid models designed for on-device deployment. It builds on the LFM2 architecture with a 131,072-token context window and agentic post-training. Best-in-class agent: competitive with models 4x larger on tool use, instruction following and multi-step agentic tasks.
Architecturally it is a hybrid of 22 double-gated convolution blocks and 8 Grouped-Query Attention (GQA) full-attention layers (30 layers total, 2.69B parameters, RoPE base 10M, vocabulary 128,000). Pre-training used ~34 trillion tokens; a mid-training phase extended the context window to 128K; post-training turned the base model into an agent in four stages - supervised fine-tuning (two rounds), per-domain teacher specialization, multi-domain on-policy distillation, and agentic reinforcement learning that trains the model directly inside popular agentic harnesses (their tools, system prompts and interaction patterns).
It is text-only, speaks 16 languages (English, Arabic, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Spanish, Vietnamese, Thai, Indonesian, Hindi, Russian, Polish), uses a ChatML-like chat template with Pythonic function calling, and runs efficiently on edge hardware: 220 tok/s on an Apple M5 Max and 113 tok/s on an AMD Ryzen CPU in under 2.5 GB of memory. A 328M speculative decoding drafter (LFM2.5-2.6B-DSpark) pairs with it for ~2.6x faster decoding with identical outputs.
Recommended for agentic workloads, tool use, data extraction, RAG and long-context workflows; not recommended for agentic coding and knowledge-heavy tasks. Released August 4, 2026 under the LFM Open License v1.0.
Training Data 34 trillion tokens pre-training, mid-training for 128K context extension, post-training with SFT, per-domain teacher specialization, on-policy distillation, agentic RL.
Benchmark Scores
| Benchmark | Score | Date |
|---|---|---|
|
MMLU-Pro
knowledge
|
16.77%
|
— |
|
BFCL-V4
general_agent
|
73.57%
|
— |
|
MMLU-Redux
knowledge
|
34.45%
|
— |
|
SWE-bench Verified
coding_agent
|
6.42%
|
— |
|
Humanity's Last Exam
stem_reasoning
|
7.77%
|
— |
|
SWE-bench Pro
coding_agent
|
0.75%
|
— |
|
Terminal Bench 2.1
coding_agent
|
4.55%
|
— |
|
GPQA Diamond
stem_reasoning
|
27.21%
|
— |
|
BrowseComp-zh
general_agent
|
9.76%
|
— |
|
SuperGPQA
knowledge
|
26.20
|
— |
|
BrowseComp
general_agent
|
12.03%
|
— |
|
AA-LCR
long_context
|
6.62%
|
— |
|
Gaia2
general_agent
|
36.98%
|
— |
|
NoLiMa
long_context
|
0.30%
|
— |
|
LongBenchPro
long_context
|
30.88%
|
— |
|
GDPVal-AA v2
general_agent
|
0.25%
|
— |
|
LongBench v2
long_context
|
12.22%
|
— |
|
Claw-Eval Avg
coding_agent
|
20.83%
|
— |
|
WildClawBench
coding_agent
|
8.38%
|
— |
|
TAU2-Bench
general_agent
|
91.73%
|
— |
|
QwenClawBench
coding_agent
|
30.27%
|
— |
|
LiveCodeBench v6
stem_reasoning
|
29.88%
|
— |
|
LCB-Pro 25Q2 (Easy)
stem_reasoning
|
35.70%
|
— |
|
LCB-Pro 25Q2 (Medium)
stem_reasoning
|
0.00
|
— |
|
OJBench
stem_reasoning
|
12.63%
|
— |
|
SciCode
|
24.26%
|
— |
|
AIME 2025
stem_reasoning
|
19.55%
|
— |
|
AIME 26
stem_reasoning
|
31.12%
|
— |
|
HMMT Feb 26
stem_reasoning
|
17.10%
|
— |
|
MATH-500
math
|
30.88%
|
— |
|
IFEval
instruction_following
|
97.34%
|
— |
|
Multi-IF
instruction_following
|
97.33%
|
— |
|
TAU3-Bench
general_agent
|
8.86%
|
— |
|
IFBench
instruction_following
|
60.40%
|
— |
Model Tree, Spaces, Collection and Paper
Model tree for LiquidAI/LFM2.5-2.6B
Base model
Finetuned
(11)
this model
Adapters
Finetunes
Merges
Quantizations
Spaces using LiquidAI/LFM2.5-2.6B 11
Collection including LiquidAI/LFM2.5-2.6B
[
💧 LFM2.5
Collection
Collection of post-trained and base LFM2.5 models. • 16 items • Updated Aug 4 • 220
](https://huggingface.co/collections/LiquidAI/lfm25)
Paper for LiquidAI/LFM2.5-2.6B
[
LFM2 Technical Report
Paper • 2511.23404 • Published Nov 28, 2025 • 72
](https://huggingface.co/papers/2511.23404)
Article mentioning LiquidAI/LFM2.5-2.6B
[
Deploy local agents everywhere with LFM2.5-2.6B
LiquidAI
•
Aug 4
• 96
](https://huggingface.co/blog/LiquidAI/lfm2-5-2-6b)
Citations
Citation
@article{liquidAI202626B,
author = {Liquid AI},
title = {LFM2.5-2.6B: Agents Everywhere},
journal = {Liquid AI Blog},
year = {2026},
note = {www.liquid.ai/blog/lfm2-5-2-6b},
}
@article{liquidai2025lfm2,
title = {LFM2 Technical Report},
author = {Liquid AI},
journal = {arXiv preprint arXiv:2511.23404},
year = {2025}
}
Model size
3B params
Tensor type
BF16
·
CPU and GPU Inference Throughput
CPU Inference
Due to its efficient LFM2 architecture, LFM2.5-2.6B is the fastest model we tested, with decode speeds of 220 tokens/s on an M5 Max and 113 tokens/s on a Ryzen AI Max+ 395. At 30 tokens/s, it allows you to run capable agents even on a phone.
GPU Inference
LFM2.5-2.6B is the fastest model in its size class, reaching almost 15K output tokens per second at high concurrency, roughly 1.3B tokens per day on a single H100.
Benchmarks vs sub-10B models
Benchmarks
We compared LFM2.5-2.6B with relevant sub-10B models on a diverse suite of benchmarks.
| Benchmark | LFM2.5-2.6B (2.6B) | gemma-4-E2B-it (5.1B) | gemma-4-E4B-it (8B) | Qwen3.5-4B (4.7B) | Qwen3.5-9B (9.7B) |
|---|---|---|---|---|---|
| AA-Omni-Public Index | -29.50 | -74.47 | -49.03 | -54.30 | -50.43 |
| AA-Omni-Public Acc | 8.13 | 6.37 | 8.33 | 17.63 | 21.30 |
| AA-Omni-Public Non-hallu | 59.04 | 13.67 | 37.42 | 12.66 | 8.84 |
| AIME25 | 51.87 | 26.33 | 34.27 | 49.33 | 56.07 |
| LiveCodeBenchv6 | 59.41 | 54.92 | 63.77 | 60.85 | 69.86 |
| IFBench | 59.17 | 34.08 | 39.24 | 48.40 | 56.47 |
| Multi-IF | 80.07 | 69.44 | 77.35 | 55.67 | 62.55 |
| IFStruct | 85.49 | 64.85 | 76.65 | 36.25 | 78.50 |
| BFCLv4 | 56.88 | 36.98 | 46.39 | 50.56 | 60.13 |
| ToolSandbox | 77.83 | 52.40 | 65.00 | 75.55 | 76.44 |
| τ³-Bench Banking | 5.67 | 3.35 | 4.12 | 5.45 | 5.15 |
| Claw-Eval average (EN) | 62.85 | 53.14 | 58.02 | 62.28 | 66.53 |
| PinchBench | 68.22 | 44.24 | 55.09 | 71.26 | 71.45 |
| BrowseComp+ (OpenClaw) | 26.89 | 8.31 | 15.90 | 24.46 | 27.23 |
Fine-Tuning Recipes
🔧 Fine-Tuning
We recommend fine-tuning LFM2.5 for your specific use case to achieve the best results.
| Name | Description | Docs | Notebook |
|---|---|---|---|
| CPT (Unsloth) | Continued Pre-Training using Unsloth for text completion. | Link | ![]() |
| CPT (Unsloth) | Continued Pre-Training using Unsloth for translation. | Link | ![]() |
| SFT (Unsloth) | Supervised Fine-Tuning with LoRA using Unsloth. | Link | ![]() |
| SFT (TRL) | Supervised Fine-Tuning with LoRA using TRL. | Link | ![]() |
| DPO (TRL) | Direct Preference Optimization with LoRA using TRL. | Link | ![]() |
| GRPO (TRL) | GRPO with LoRA using TRL. | Link | ![]() |
How to Use: Quick Start and Agent Use
How to use
LFM2.5-2.6B can be used for direct inference or as a backend for agentic workflows.
Quick start
Get started with Transformers (compatible with transformers>=5.0.0):
from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer
model_id = "LiquidAI/LFM2.5-2.6B"
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
dtype="bfloat16",
## attn_implementation="flash_attention_2" <- uncomment on compatible GPU
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
streamer = TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)
prompt = "What is C. elegans?"
input_ids = tokenizer.apply_chat_template(
[{"role": "user", "content": prompt}],
add_generation_prompt=True,
return_tensors="pt",
tokenize=True,
)["input_ids"].to(model.device)
output = model.generate(
input_ids,
do_sample=True,
temperature=0.1,
top_k=50,
repetition_penalty=1.1,
max_new_tokens=512,
streamer=streamer,
)
Agent Use
LFM2.5-2.6B supports tool calling for agentic workflows. Serve it locally with any OpenAI-compatible backend (see 🏃 Inference, then configure your agent harness to connect to it. For full setup instructions including installation and additional options, see our Agent Harnesses guide.
Note: The port depends on your serving backend — llama.cpp and MLX use 8080, vLLM uses 8000, SGLang uses 30000, and LM Studio uses 1234. Adjust the URLs below accordingly.
Hermes
Either use the interactive wizard or set it directly:
hermes config set model.provider custom
hermes config set model.base_url http://localhost:8080/v1
hermes config set model.default LFM2.5-2.6B
hermes config set model.context_length 131072
hermes config set model.api_mode chat_completions
hermes config set agent.tool_use_enforcement true
OpenClaw
Add to your config to models.providers:
local: {
baseUrl: "http://localhost:8080/v1",
apiKey: "sk-local",
api: "openai-completions",
models: [{
id: "LFM2.5-2.6B",
name: "LFM2.5-2.6B",
contextWindow: 131072,
maxTokens: 8192,
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 }
}]
}
Pi
Add to your config to ~/.pi/agent/models.json:
{
"providers": {
"local": {
"baseUrl": "http://localhost:8080/v1",
"api": "openai-completions",
"apiKey": "local",
"models": [{ "id": "LFM2.5-2.6B" }]
}
}
}
Inference Frameworks (Transformers, vLLM, SGLang, llama.cpp, MLX)
🏃 Inference
LFM2.5 is supported by many inference frameworks. See the Inference documentation for the full list.
| Name | Description | Docs | Notebook |
|---|---|---|---|
| Transformers | Simple inference with direct access to model internals. | Link | ![]() |
| vLLM | High-throughput production deployments with GPU. | Link | ![]() |
| SGLang | High-throughput production deployments with GPU. | Link | — |
| llama.cpp | Cross-platform inference with CPU offloading. | Link | ![]() |
| MLX | Apple's machine learning framework optimized for Apple Silicon. | Link | — |
| LM Studio | Desktop application for running LLMs locally. | Link | — |
⚡ Faster decoding: attach LFM2.5-2.6B-DSpark, a 328M speculative-decoding drafter, for ~2.6x faster decoding in SGLang and on Apple silicon via Metal with exactly the same outputs.
Training Pipeline (mid-training, SFT, teacher specialization, distillation, agentic RL)
Training
LFM2.5-2.6B is pre-trained on ~34T tokens, with a mid-training phase that extends the context window to 128K. Post-training then turns the base model into an agent in four stages: supervised fine-tuning (two rounds), per-domain teacher specialization, multi-domain on-policy distillation, and agentic reinforcement learning.
In particular, agentic reinforcement learning allows us to directly train the model inside popular agentic harnesses. It exposes the model to their tools, system prompts, and interaction patterns, helping it work reliably across agent environments.
Tool Use: Function Calling
Tool Use
LFM2.5 supports function calling in four steps:
- Function definition: Provide the list of tools as a JSON object in the system prompt, or use
tokenizer.apply_chat_template()withtools=.... - Function call: By default, LFM2.5 writes Pythonic function calls (a Python list between
<|tool_call_start|>and<|tool_call_end|>special tokens), as the assistant answer. You can override this behavior by asking the model to output JSON function calls in the system prompt. - Function execution: Execute the call and return the result with the
toolrole. - Final answer: LFM2.5 interprets the tool output and returns a plain-text answer addressing the original prompt.
See the Tool Use documentation for the full guide. Example:
<|startoftext|><|im_start|>system
List of tools: [{"name": "get_candidate_status", "description": "Retrieves the current status of a candidate in the recruitment process", "parameters": {"type": "object", "properties": {"candidate_id": {"type": "string", "description": "Unique identifier for the candidate"}}, "required": ["candidate_id"]}}]<|im_end|>
<|im_start|>user
What is the current status of candidate ID 12345?<|im_end|>
<|im_start|>assistant
<|tool_call_start|>[get_candidate_status(candidate_id="12345")]<|tool_call_end|>Checking the current status of candidate ID 12345.<|im_end|>
<|im_start|>tool
[{"candidate_id": "12345", "status": "Interview Scheduled", "position": "Clinical Research Associate", "date": "2023-11-20"}]<|im_end|>
<|im_start|>assistant
The candidate with ID 12345 is currently in the "Interview Scheduled" stage for the position of Clinical Research Associate, with an interview date set for 2023-11-20.<|im_end|>
Chat Template (ChatML-like)
Chat Template
LFM2.5 uses a ChatML-like format. See the Chat Template documentation for details. Example:
<|startoftext|><|im_start|>system
You are a helpful assistant trained by Liquid AI.<|im_end|>
<|im_start|>user
What is C. elegans?<|im_end|>
<|im_start|>assistant
You can use tokenizer.apply_chat_template() to format your messages automatically.
💡 Note: LFM2.5-2.6B is a pure reasoning model that always thinks before it answers. It adds a
<think>tag directly in the chat template when starting an assistant answer.
Model Variants and Formats (GGUF, ONNX, MLX, DSpark)
cpp and compatible tools. Optimized for CPU inference and local deployment with reduced memory usage. | | LFM2.5-2.6B-ONNX | ONNX Runtime format for cross-platform deployment. Enables hardware-accelerated inference across diverse environments (cloud, edge, mobile). | | LFM2.5-2.6B-MLX | MLX format for Apple Silicon. Optimized for fast inference on Mac devices using the MLX framework. | | LFM2.5-2.6B-DSpark | Speculative decoding drafter (328M). Pair it with this model for ~2.6x faster decoding with identical outputs. |
We recommend using it for agentic workloads, tool use, data extraction, RAG, and long-context workflows. It is not recommended for agentic coding and knowledge-heavy tasks.
Model Details: Parameters, Layers, Context, Languages, Generation Parameters
🗒️ Model Details
| Model | Parameters | Description |
|---|---|---|
| LFM2.5-2.6B-Base | 2.6B | Pre-trained base model for fine-tuning |
| LFM2.5-2.6B | 2.6B | Post-trained for agentic workloads |
LFM2.5-2.6B is a general-purpose text-only model with the following features:
- Total parameters: 2.69B
- Number of layers: 30 (22 double-gated short convolution blocks + 8 GQA)
- Training budget: 34 trillion tokens
- Vocabulary size: 128,000
- Context length: 131,072 tokens
- Languages: English, Arabic, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Spanish, Vietnamese, Thai, Indonesian, Hindi, Russian, Polish
- Generation parameters:
temperature: 0.1top_k: 50repetition_penalty: 1.1
| Model | Description |
|---|---|
| LFM2.5-2.6B | Original model checkpoint in native format. Best for fine-tuning or inference with Transformers, vLLM, and SGLang. |
| LFM2.5-2.6B-GGUF | Quantized format for llama. |
LFM2.5-2.6B: Best-in-Class On-Device Agent Overview
LFM2.5-2.6B
LFM2.5-2.6B is part of LFM2.5, a family of hybrid models designed for on-device deployment. It builds on the LFM2 architecture with a 128K context window and agentic post-training.
- Best-in-class agent: Competitive with models 4x larger on tool use, instruction following, and multi-step agentic tasks.
- Agentic reinforcement learning: Trained inside the most popular agentic harnesses to improve compatibility.
- Efficient inference: 220 tok/s on an Apple M5 Max and 113 tok/s on an AMD Ryzen CPU, in under 2.5 GB of memory.
Find more information about LFM2.5-2.6B in our blog post.
💻 Demos: Try LFM2.5-2.6B's agentic capabilities in a Hugging Face space without any setup: Research Agent in your browser: helps you research a specific question and generates a summary
Architecture
- Attention
- Hybrid Attention (32:8)
- Layers
- 30
- Hidden size
- 2048
- Context
- 131K tokens
- Parameters
- 2690M
Source: Hugging Face config.json · Lfm2ForCausalLM · exact layer pattern · model repo
LFM2.5-2.6B-DSpark (328M, ~2.6x faster decoding)
Training Pipeline
-
1
pretraining
Pre-training (~34T tokens)
Pre-trained on ~34 trillion tokens on the LFM2 hybrid architecture.
-
2
other
Mid-training: 128K context extension
Mid-training phase extending the context window to 128K (131,072 tokens).
-
3
sft
Supervised fine-tuning (two rounds)
First post-training stage turning the base model toward agentic workloads.
-
4
other
Per-domain teacher specialization
Teacher models specialized per domain.
-
5
other
Multi-domain on-policy distillation
On-policy distillation from the specialized teachers into the student.
-
6
rl
Agentic reinforcement learning
RL inside popular agentic harnesses: exposes the model to their tools, system prompts and interaction patterns for reliable cross-environment agent behavior.
Training & Evaluation Datasets
| Name | Role | Size | Modalities | Collection |
|---|---|---|---|---|
| AA-Omni-Public (Artificial Analysis) | evaluation | — | — | |
| IFBench / Multi-IF / IFStruct | evaluation | — | — | |
| AIME25 / LiveCodeBench v6 | evaluation | — | — | |
| BFCL v4 / ToolSandbox / Tau-cubed-Bench Banking | evaluation | — | — | |
| Claw-Eval / PinchBench / BrowseComp+ (OpenClaw) | evaluation | — | — |
Linked Resources
LFM2.5-2.6B: Deploy Agents Everywhere
https://www.liquid.ai/blog/lfm2-5-2-6b
LFM2 Technical Report
https://arxiv.org/abs/2511.23404
Liquid LFM documentation
https://docs.liquid.ai/lfm/getting-started/welcome
LFM2.5-2.6B Research Agent (WebGPU, in-browser)
https://huggingface.co/spaces/LiquidAI/LFM2.5-2.6B-WebGPU
BibTeX citations (Liquid AI, 2026)
https://huggingface.co/LiquidAI/LFM2.5-2.6B





