LFM2.5-2.6B

LiquidAI

Parameters

2.7B

Architecture

LFM2 (Hybrid)

Released

04.08.2026

License

LFM Open License v1.0

Open Weights Commercial Use Multimodal BF16 LFM en ar zh fr de it ja ko pt es vi th id hi ru pl

Input Modalities

text

Output Modalities

text

Context (native)

131,072 tokens

Context (extended)

131,072 tokens

Openness Index Score 70.0/100

About

LFM2.5-2.6B is Liquid AI's post-trained agent model - part of the LFM2.5 family of hybrid models designed for on-device deployment. It builds on the LFM2 architecture with a 131,072-token context window and agentic post-training. Best-in-class agent: competitive with models 4x larger on tool use, instruction following and multi-step agentic tasks.

Architecturally it is a hybrid of 22 double-gated convolution blocks and 8 Grouped-Query Attention (GQA) full-attention layers (30 layers total, 2.69B parameters, RoPE base 10M, vocabulary 128,000). Pre-training used ~34 trillion tokens; a mid-training phase extended the context window to 128K; post-training turned the base model into an agent in four stages - supervised fine-tuning (two rounds), per-domain teacher specialization, multi-domain on-policy distillation, and agentic reinforcement learning that trains the model directly inside popular agentic harnesses (their tools, system prompts and interaction patterns).

It is text-only, speaks 16 languages (English, Arabic, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Spanish, Vietnamese, Thai, Indonesian, Hindi, Russian, Polish), uses a ChatML-like chat template with Pythonic function calling, and runs efficiently on edge hardware: 220 tok/s on an Apple M5 Max and 113 tok/s on an AMD Ryzen CPU in under 2.5 GB of memory. A 328M speculative decoding drafter (LFM2.5-2.6B-DSpark) pairs with it for ~2.6x faster decoding with identical outputs.

Recommended for agentic workloads, tool use, data extraction, RAG and long-context workflows; not recommended for agentic coding and knowledge-heavy tasks. Released August 4, 2026 under the LFM Open License v1.0.

Training Data 34 trillion tokens pre-training, mid-training for 128K context extension, post-training with SFT, per-domain teacher specialization, on-policy distillation, agentic RL.

Benchmark Scores

Benchmark Score Date
MMLU-Pro
knowledge
16.77%
—
BFCL-V4
general_agent
73.57%
—
MMLU-Redux
knowledge
34.45%
—
SWE-bench Verified
coding_agent
6.42%
—
Humanity's Last Exam
stem_reasoning
7.77%
—
SWE-bench Pro
coding_agent
0.75%
—
Terminal Bench 2.1
coding_agent
4.55%
—
GPQA Diamond
stem_reasoning
27.21%
—
BrowseComp-zh
general_agent
9.76%
—
SuperGPQA
knowledge
26.20
—
BrowseComp
general_agent
12.03%
—
AA-LCR
long_context
6.62%
—
Gaia2
general_agent
36.98%
—
NoLiMa
long_context
0.30%
—
LongBenchPro
long_context
30.88%
—
GDPVal-AA v2
general_agent
0.25%
—
LongBench v2
long_context
12.22%
—
Claw-Eval Avg
coding_agent
20.83%
—
WildClawBench
coding_agent
8.38%
—
TAU2-Bench
general_agent
91.73%
—
QwenClawBench
coding_agent
30.27%
—
LiveCodeBench v6
stem_reasoning
29.88%
—
LCB-Pro 25Q2 (Easy)
stem_reasoning
35.70%
—
LCB-Pro 25Q2 (Medium)
stem_reasoning
0.00
—
OJBench
stem_reasoning
12.63%
—
SciCode
24.26%
—
AIME 2025
stem_reasoning
19.55%
—
AIME 26
stem_reasoning
31.12%
—
HMMT Feb 26
stem_reasoning
17.10%
—
MATH-500
math
30.88%
—
IFEval
instruction_following
97.34%
—
Multi-IF
instruction_following
97.33%
—
TAU3-Bench
general_agent
8.86%
—
IFBench
instruction_following
60.40%
—

Model Tree, Spaces, Collection and Paper

Model tree for LiquidAI/LFM2.5-2.6B

Base model

LiquidAI/LFM2.5-2.6B-Base

Finetuned

(11)

this model

Adapters

19 models

Finetunes

39 models

Merges

1 model

Quantizations

91 models

Spaces using LiquidAI/LFM2.5-2.6B 11

Collection including LiquidAI/LFM2.5-2.6B

[

💧 LFM2.5

Collection

Collection of post-trained and base LFM2.5 models. • 16 items • Updated Aug 4 • 220

](https://huggingface.co/collections/LiquidAI/lfm25)

Paper for LiquidAI/LFM2.5-2.6B

[

LFM2 Technical Report

Paper • 2511.23404 • Published Nov 28, 2025 • 72

](https://huggingface.co/papers/2511.23404)

Article mentioning LiquidAI/LFM2.5-2.6B

[

Deploy local agents everywhere with LFM2.5-2.6B

LiquidAI

•

Aug 4

• 96

](https://huggingface.co/blog/LiquidAI/lfm2-5-2-6b)

Citations

Citation

@article{liquidAI202626B,
  author  = {Liquid AI},
  title   = {LFM2.5-2.6B: Agents Everywhere},
  journal = {Liquid AI Blog},
  year    = {2026},
  note    = {www.liquid.ai/blog/lfm2-5-2-6b},
}

@article{liquidai2025lfm2,
  title   = {LFM2 Technical Report},
  author  = {Liquid AI},
  journal = {arXiv preprint arXiv:2511.23404},
  year    = {2025}
}

Safetensors

Model size

3B params

Tensor type

BF16

·

CPU and GPU Inference Throughput

CPU Inference

Due to its efficient LFM2 architecture, LFM2.5-2.6B is the fastest model we tested, with decode speeds of 220 tokens/s on an M5 Max and 113 tokens/s on a Ryzen AI Max+ 395. At 30 tokens/s, it allows you to run capable agents even on a phone.

GPU Inference

LFM2.5-2.6B is the fastest model in its size class, reaching almost 15K output tokens per second at high concurrency, roughly 1.3B tokens per day on a single H100.

Benchmarks vs sub-10B models

Benchmarks

We compared LFM2.5-2.6B with relevant sub-10B models on a diverse suite of benchmarks.

Benchmark LFM2.5-2.6B (2.6B) gemma-4-E2B-it (5.1B) gemma-4-E4B-it (8B) Qwen3.5-4B (4.7B) Qwen3.5-9B (9.7B)
AA-Omni-Public Index -29.50 -74.47 -49.03 -54.30 -50.43
AA-Omni-Public Acc 8.13 6.37 8.33 17.63 21.30
AA-Omni-Public Non-hallu 59.04 13.67 37.42 12.66 8.84
AIME25 51.87 26.33 34.27 49.33 56.07
LiveCodeBenchv6 59.41 54.92 63.77 60.85 69.86
IFBench 59.17 34.08 39.24 48.40 56.47
Multi-IF 80.07 69.44 77.35 55.67 62.55
IFStruct 85.49 64.85 76.65 36.25 78.50
BFCLv4 56.88 36.98 46.39 50.56 60.13
ToolSandbox 77.83 52.40 65.00 75.55 76.44
τ³-Bench Banking 5.67 3.35 4.12 5.45 5.15
Claw-Eval average (EN) 62.85 53.14 58.02 62.28 66.53
PinchBench 68.22 44.24 55.09 71.26 71.45
BrowseComp+ (OpenClaw) 26.89 8.31 15.90 24.46 27.23

Fine-Tuning Recipes

🔧 Fine-Tuning

We recommend fine-tuning LFM2.5 for your specific use case to achieve the best results.

Name Description Docs Notebook
CPT (Unsloth) Continued Pre-Training using Unsloth for text completion. Link Colab link
CPT (Unsloth) Continued Pre-Training using Unsloth for translation. Link Colab link
SFT (Unsloth) Supervised Fine-Tuning with LoRA using Unsloth. Link Colab link
SFT (TRL) Supervised Fine-Tuning with LoRA using TRL. Link Colab link
DPO (TRL) Direct Preference Optimization with LoRA using TRL. Link Colab link
GRPO (TRL) GRPO with LoRA using TRL. Link Colab link

How to Use: Quick Start and Agent Use

How to use

LFM2.5-2.6B can be used for direct inference or as a backend for agentic workflows.

Quick start

Get started with Transformers (compatible with transformers>=5.0.0):

from transformers import AutoModelForCausalLM, AutoTokenizer, TextStreamer

model_id = "LiquidAI/LFM2.5-2.6B"
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    device_map="auto",
    dtype="bfloat16",
##   attn_implementation="flash_attention_2" <- uncomment on compatible GPU
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
streamer = TextStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)

prompt = "What is C. elegans?"

input_ids = tokenizer.apply_chat_template(
    [{"role": "user", "content": prompt}],
    add_generation_prompt=True,
    return_tensors="pt",
    tokenize=True,
)["input_ids"].to(model.device)

output = model.generate(
    input_ids,
    do_sample=True,
    temperature=0.1,
    top_k=50,
    repetition_penalty=1.1,
    max_new_tokens=512,
    streamer=streamer,
)

Agent Use

LFM2.5-2.6B supports tool calling for agentic workflows. Serve it locally with any OpenAI-compatible backend (see 🏃 Inference, then configure your agent harness to connect to it. For full setup instructions including installation and additional options, see our Agent Harnesses guide.

Note: The port depends on your serving backend — llama.cpp and MLX use 8080, vLLM uses 8000, SGLang uses 30000, and LM Studio uses 1234. Adjust the URLs below accordingly.

Hermes

Either use the interactive wizard or set it directly:

hermes config set model.provider custom
hermes config set model.base_url http://localhost:8080/v1
hermes config set model.default LFM2.5-2.6B
hermes config set model.context_length 131072
hermes config set model.api_mode chat_completions
hermes config set agent.tool_use_enforcement true

OpenClaw

Add to your config to models.providers:

local: {
  baseUrl: "http://localhost:8080/v1",
  apiKey: "sk-local",
  api: "openai-completions",
  models: [{
    id: "LFM2.5-2.6B",
    name: "LFM2.5-2.6B",
    contextWindow: 131072,
    maxTokens: 8192,
    cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 }
  }]
}

Pi

Add to your config to ~/.pi/agent/models.json:

{
  "providers": {
    "local": {
      "baseUrl": "http://localhost:8080/v1",
      "api": "openai-completions",
      "apiKey": "local",
      "models": [{ "id": "LFM2.5-2.6B" }]
    }
  }
}

Inference Frameworks (Transformers, vLLM, SGLang, llama.cpp, MLX)

🏃 Inference

LFM2.5 is supported by many inference frameworks. See the Inference documentation for the full list.

Name Description Docs Notebook
Transformers Simple inference with direct access to model internals. Link Colab link
vLLM High-throughput production deployments with GPU. Link Colab link
SGLang High-throughput production deployments with GPU. Link —
llama.cpp Cross-platform inference with CPU offloading. Link Colab link
MLX Apple's machine learning framework optimized for Apple Silicon. Link —
LM Studio Desktop application for running LLMs locally. Link —

⚡ Faster decoding: attach LFM2.5-2.6B-DSpark, a 328M speculative-decoding drafter, for ~2.6x faster decoding in SGLang and on Apple silicon via Metal with exactly the same outputs.

Training Pipeline (mid-training, SFT, teacher specialization, distillation, agentic RL)

Training

LFM2.5-2.6B is pre-trained on ~34T tokens, with a mid-training phase that extends the context window to 128K. Post-training then turns the base model into an agent in four stages: supervised fine-tuning (two rounds), per-domain teacher specialization, multi-domain on-policy distillation, and agentic reinforcement learning.

In particular, agentic reinforcement learning allows us to directly train the model inside popular agentic harnesses. It exposes the model to their tools, system prompts, and interaction patterns, helping it work reliably across agent environments.

Tool Use: Function Calling

Tool Use

LFM2.5 supports function calling in four steps:

  1. Function definition: Provide the list of tools as a JSON object in the system prompt, or use tokenizer.apply_chat_template() with tools=....
  2. Function call: By default, LFM2.5 writes Pythonic function calls (a Python list between <|tool_call_start|> and <|tool_call_end|> special tokens), as the assistant answer. You can override this behavior by asking the model to output JSON function calls in the system prompt.
  3. Function execution: Execute the call and return the result with the tool role.
  4. Final answer: LFM2.5 interprets the tool output and returns a plain-text answer addressing the original prompt.

See the Tool Use documentation for the full guide. Example:

<|startoftext|><|im_start|>system
List of tools: [{"name": "get_candidate_status", "description": "Retrieves the current status of a candidate in the recruitment process", "parameters": {"type": "object", "properties": {"candidate_id": {"type": "string", "description": "Unique identifier for the candidate"}}, "required": ["candidate_id"]}}]<|im_end|>
<|im_start|>user
What is the current status of candidate ID 12345?<|im_end|>
<|im_start|>assistant
<|tool_call_start|>[get_candidate_status(candidate_id="12345")]<|tool_call_end|>Checking the current status of candidate ID 12345.<|im_end|>
<|im_start|>tool
[{"candidate_id": "12345", "status": "Interview Scheduled", "position": "Clinical Research Associate", "date": "2023-11-20"}]<|im_end|>
<|im_start|>assistant
The candidate with ID 12345 is currently in the "Interview Scheduled" stage for the position of Clinical Research Associate, with an interview date set for 2023-11-20.<|im_end|>

Chat Template (ChatML-like)

Chat Template

LFM2.5 uses a ChatML-like format. See the Chat Template documentation for details. Example:

<|startoftext|><|im_start|>system
You are a helpful assistant trained by Liquid AI.<|im_end|>
<|im_start|>user
What is C. elegans?<|im_end|>
<|im_start|>assistant

You can use tokenizer.apply_chat_template() to format your messages automatically.

💡 Note: LFM2.5-2.6B is a pure reasoning model that always thinks before it answers. It adds a <think> tag directly in the chat template when starting an assistant answer.

Model Variants and Formats (GGUF, ONNX, MLX, DSpark)

cpp and compatible tools. Optimized for CPU inference and local deployment with reduced memory usage. | | LFM2.5-2.6B-ONNX | ONNX Runtime format for cross-platform deployment. Enables hardware-accelerated inference across diverse environments (cloud, edge, mobile). | | LFM2.5-2.6B-MLX | MLX format for Apple Silicon. Optimized for fast inference on Mac devices using the MLX framework. | | LFM2.5-2.6B-DSpark | Speculative decoding drafter (328M). Pair it with this model for ~2.6x faster decoding with identical outputs. |

We recommend using it for agentic workloads, tool use, data extraction, RAG, and long-context workflows. It is not recommended for agentic coding and knowledge-heavy tasks.

Model Details: Parameters, Layers, Context, Languages, Generation Parameters

🗒️ Model Details

Model Parameters Description
LFM2.5-2.6B-Base 2.6B Pre-trained base model for fine-tuning
LFM2.5-2.6B 2.6B Post-trained for agentic workloads

LFM2.5-2.6B is a general-purpose text-only model with the following features:

  • Total parameters: 2.69B
  • Number of layers: 30 (22 double-gated short convolution blocks + 8 GQA)
  • Training budget: 34 trillion tokens
  • Vocabulary size: 128,000
  • Context length: 131,072 tokens
  • Languages: English, Arabic, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Spanish, Vietnamese, Thai, Indonesian, Hindi, Russian, Polish
  • Generation parameters:
    • temperature: 0.1
    • top_k: 50
    • repetition_penalty: 1.1
Model Description
LFM2.5-2.6B Original model checkpoint in native format. Best for fine-tuning or inference with Transformers, vLLM, and SGLang.
LFM2.5-2.6B-GGUF Quantized format for llama.

LFM2.5-2.6B: Best-in-Class On-Device Agent Overview

LFM2.5-2.6B

LFM2.5-2.6B is part of LFM2.5, a family of hybrid models designed for on-device deployment. It builds on the LFM2 architecture with a 128K context window and agentic post-training.

  • Best-in-class agent: Competitive with models 4x larger on tool use, instruction following, and multi-step agentic tasks.
  • Agentic reinforcement learning: Trained inside the most popular agentic harnesses to improve compatibility.
  • Efficient inference: 220 tok/s on an Apple M5 Max and 113 tok/s on an AMD Ryzen CPU, in under 2.5 GB of memory.

Find more information about LFM2.5-2.6B in our blog post.

💻 Demos: Try LFM2.5-2.6B's agentic capabilities in a Hugging Face space without any setup: Research Agent in your browser: helps you research a specific question and generates a summary

Architecture

Decoder Block ×30 input Embedding vocab 128K · d 2048 Full Attention Hybrid 32:8 ×8 Dense FFN SwiGLU · d 11K Final RMSNorm LM Head tied with embedding output
Attention
Hybrid Attention (32:8)
Layers
30
Hidden size
2048
Context
131K tokens
Parameters
2690M

Source: Hugging Face config.json · Lfm2ForCausalLM · exact layer pattern · model repo

Type: Hybrid dense: 22 double-gated short-conv blocks + 8 GQA full-attention layers (Lfm2ForCausalLM)
Attention: GQA (32 query / 8 KV heads) on full-attention layers; RoPE theta 10M; double-gated conv blocks (L cache 3)
Decoder: Dense decoder-only hybrid (conv + attention)
Layers 30
Total parameters 2690M
Context length 131K
Extended context 131K
Attention heads 32
KV heads 8
Hidden size 2048
Vocabulary 128K
FFN dim 11K
Precision bfloat16
RoPE θ 10M
Attention Layers 8
Chat Template ChatML-like
Conv Layers 22
Tied embeddings Yes
Training tokens 34000000M
Speculative Drafter

LFM2.5-2.6B-DSpark (328M, ~2.6x faster decoding)

Training Pipeline

  1. 1
    pretraining

    Pre-training (~34T tokens)

    Pre-trained on ~34 trillion tokens on the LFM2 hybrid architecture.

  2. 2
    other

    Mid-training: 128K context extension

    Mid-training phase extending the context window to 128K (131,072 tokens).

  3. 3
    sft

    Supervised fine-tuning (two rounds)

    First post-training stage turning the base model toward agentic workloads.

  4. 4
    other

    Per-domain teacher specialization

    Teacher models specialized per domain.

  5. 5
    other

    Multi-domain on-policy distillation

    On-policy distillation from the specialized teachers into the student.

  6. 6
    rl

    Agentic reinforcement learning

    RL inside popular agentic harnesses: exposes the model to their tools, system prompts and interaction patterns for reliable cross-environment agent behavior.

Training & Evaluation Datasets

NameRoleSizeModalitiesCollection
AA-Omni-Public (Artificial Analysis) evaluation — —
IFBench / Multi-IF / IFStruct evaluation — —
AIME25 / LiveCodeBench v6 evaluation — —
BFCL v4 / ToolSandbox / Tau-cubed-Bench Banking evaluation — —
Claw-Eval / PinchBench / BrowseComp+ (OpenClaw) evaluation — —

Related Models