MiMo-V2.5 is a native, fully multimodal large model designed for agent scenarios; it can see, hear, and read, and translate understanding into action. It has over 1 trillion total parameters (42B active parameters), employs an innovative hybrid-attention architecture, and supports an ultra-long 1M context length. Built on a powerful model base, we continuously scale compute across broader agent scenarios, further expanding the agent’s action space and achieving an important generalization from coding to claw.
This model will be retired on 2026-10-21. Plan to migrate to mimo-v2.6-flash.
Calls still work until then. All model retirements
Pricing
Input Modalities
- Text
- Vision
- Audio
- Video
Output Modalities
- Text
Context length
- 1.05M tokens
Max output
- 131K tokens
Capabilities
- Thinking
- Streaming
- Tool calling
- Web search
- URL context
- Code interpreter
- Computer use
- File search
- Memory tool
- Structured outputs
- Structured decision
- Citations
- Prompt caching
- Background mode
- Server-side sessions
Providers
Xiaomi xiaomi-mimo-v2.5
Pricing$0.155$0.31
Cache Read$0.0031/M tokens
Web Search$0.005/request
Context256K
Max output0
Latency4.5S
Throughput27.1TPS
Uptime
99.94% uptime 2 days ago
100.00% uptime yesterday
99.93% uptime today
Deepinfra deepinfra-mimo-v2.5
Pricing$0.44$2.2
Cache Read$0.088/M tokens
Web Search$0.005/request
Context256K
Max output0
Latency1.7S
Throughput41.3TPS
Uptime
0.00% uptime 2 days ago
100.00% uptime yesterday
0.00% uptime today
Gmicloud gmicloud-mimo-v2.5
Pricing$0.155$0.31
Cache Read$0.0031/M tokens
Web Search$0.005/request
Context1M
Max output0
Latency-
Throughput-
Uptime
0.00% uptime 2 days ago
100.00% uptime yesterday
0.00% uptime today
Performance for mimo-v2.5
Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).
Uptime
Loading...
Latency
Loading...
Throughput
Loading...
Try this model
Python
