MiMo V2.5
Xiaomi logo

MiMo V2.5

mimo-v2.5llms.txt
Xiaomi
New
MiMo-V2.5 is a native, fully multimodal large model designed for agent scenarios; it can see, hear, and read, and translate understanding into action. It has over 1 trillion total parameters (42B active parameters), employs an innovative hybrid-attention architecture, and supports an ultra-long 1M context length. Built on a powerful model base, we continuously scale compute across broader agent scenarios, further expanding the agent’s action space and achieving an important generalization from coding to claw.
This model will be retired on 2026-10-21. Plan to migrate to mimo-v2.6-flash.
Calls still work until then. All model retirements

Pricing

PricingCache ReadWeb Search
$0.155$0.31
$0.0031/M tokens$0.005/request

Input Modalities

  • Text
  • Vision
  • Audio
  • Video

Output Modalities

  • Text

Context length

  • 1.05M tokens

Max output

  • 131K tokens

Capabilities

  • Thinking
  • Streaming
  • Tool calling
  • Web search
  • URL context
  • Code interpreter
  • Computer use
  • File search
  • Memory tool
  • Structured outputs
  • Structured decision
  • Citations
  • Prompt caching
  • Background mode
  • Server-side sessions

Providers

Xiaomi xiaomi-mimo-v2.5
Pricing$0.155$0.31
Cache Read$0.0031/M tokens
Web Search$0.005/request
Context256K
Max output0
Latency4.5S
Throughput27.1TPS
Uptime
99.94% uptime 2 days ago
100.00% uptime yesterday
99.93% uptime today
Deepinfra deepinfra-mimo-v2.5
Pricing$0.44$2.2
Cache Read$0.088/M tokens
Web Search$0.005/request
Context256K
Max output0
Latency1.7S
Throughput41.3TPS
Uptime
0.00% uptime 2 days ago
100.00% uptime yesterday
0.00% uptime today
Gmicloud gmicloud-mimo-v2.5
Pricing$0.155$0.31
Cache Read$0.0031/M tokens
Web Search$0.005/request
Context1M
Max output0
Latency-
Throughput-
Uptime
0.00% uptime 2 days ago
100.00% uptime yesterday
0.00% uptime today

Performance for mimo-v2.5

Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).

Uptime
Loading...
Latency
Loading...
Throughput
Loading...

Try this model

Python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AIHUBMIX_API_KEY"],
    base_url="https://aihubmix.com/v1",
)

response = client.chat.completions.create(
    model="mimo-v2.5",
    messages=[
      {
        "role": "user",
        "content": "Hello, how are you?"
      }
    ],
    max_tokens=1024,
    stream=False,
)

print(response.choices[0].message.content)

Frequently asked questions

What is MiMo V2.5?

MiMo-V2.5 is a native, fully multimodal large model designed for agent scenarios; it can see, hear, and read, and translate understanding into action. It has over 1 trillion total parameters (42B active parameters), employs an innovative hybrid-attention architecture, and supports an ultra-long 1M context length. Built on a powerful model base, we continuously scale compute across broader agent scenarios, further expanding the agent’s action space and achieving an important generalization from coding to claw.