
Latest Model Additions: Grok 4.7, GLM 5.3 FlashX, and MiMo-V2.6
TL;DR: We have added support for Grok 4.7, GLM 5.3 FlashX, MiMo-V2.6-Pro, and MiMo-V2.6-Flash. You can select them directly in chat, assign them to dedicated agents, or wait for automatic routing as we evaluate them for Conductor Mode.
Every AI workload has different trade-offs. Complex refactoring demands deep frontier reasoning, interactive agents need instant response loops, and large data ingestion runs require rock-bottom token pricing.
To cover these distinct profiles, we have expanded our catalog with options from xAI, Z.AI, and Xiaomi.
Here is what each model brings and where it fits in your setup.
1. Grok 4.7: Flagship Agentic Execution
Grok 4.7 is xAI's latest flagship model designed for complex coding, structured data extraction, and multi-step agent actions.
- What it delivers: Higher benchmark performance than Grok 4.6 across coding and logical problem solving, delivered at a 20% lower price point per token.
- Best workflow fit: Heavy refactoring, writing complex task scripts, parsing unstructured payloads, and deep analytical research.
- How to use it: Switch to Grok 4.7 in your main chat session for technical tasks, or set it as the primary brain for sub-agents managing code generation and technical research.
2. GLM 5.3 FlashX: Snappy Interactive Loops
GLM 5.3 FlashX from Z.AI is an ultra-fast variant engineered specifically for low-latency response cycles.
- What it delivers: Instantaneous time-to-first-token and high output velocity. It cuts down execution delay during chained tool calls.
- Best workflow fit: Fast conversational back-and-forth, status pings, triage routines, and interactive task coordination where waiting for a heavy frontier model slows you down.
- How to use it: Ideal for background monitor jobs, Telegram bots, or lightweight conversational interfaces where speed matters more than complex multi-page reasoning.
3. MiMo-V2.6-Pro: Open-Weights Multimodal Flagship
MiMo-V2.6-Pro is Xiaomi's new 1T+ parameter open-weights flagship.
- What it delivers: Native omnimodal capabilities paired with controllable reasoning effort. It processes visual inputs, documents, and code side by side while giving you direct control over how deeply the model plans before answering.
- Best workflow fit: Visual QA on design layouts, document parsing, multimodal charts, and tasks that require verified multi-step reasoning before taking action.
- How to use it: Deploy it for visual review gates, slide analysis, or complex workflows that combine vision with tool execution.
4. MiMo-V2.6-Flash: High-Throughput Multimodal Processing
MiMo-V2.6-Flash is Xiaomi's budget Mixture-of-Experts (MoE) model built for volume.
- What it delivers: Full multimodal input handling at rock-bottom pricing and high throughput.
- Best workflow fit: Processing hundreds of images or documents in bulk, continuous web scraping audits, data classification pipelines, and scheduled ingestion jobs where frontier token costs would add up quickly.
- How to use it: Perfect for background jobs running on intervals, scheduled data crawlers, or high-volume scraping tasks that need multimodal intelligence on a tight budget.
Summary: Matching Models to Tasks
| Model | Primary Focus | Best For | Key Advantage |
|---|---|---|---|
| Grok 4.7 | Flagship Reasoning | Coding, Agent Pipelines, Knowledge Work | Stronger than 4.6 at 20% lower price |
| GLM 5.3 FlashX | Low-Latency Speed | Interactive Chat, Quick Bots, Fast Loops | Minimal latency, snappy tool calling |
| MiMo-V2.6-Pro | 1T+ Multimodal | Layout QA, Visual Docs, Controlled Reasoning | Native omnimodal with tunable reasoning depth |
| MiMo-V2.6-Flash | Budget MoE Multimodal | High-Volume Scraping, Ingestion, Batch Tasks | Lowest cost per multimodal input token |
Conductor Mode Evaluation
All four models are available immediately for direct use in chat and dedicated agents. At the same time, we are evaluating them for inclusion in Conductor Mode.
Conductor Mode combines classification algorithms, task heuristics, and TypeSafe's Jev engine to match each prompt to the most efficient model based on speed, cost, and capability. Once these models pass our latency and accuracy benchmarks across agent tool-calling loops, we will roll them into Conductor's automated routing pool.
You can select any of the new models today from the model picker in chat or configure them directly in your agent settings.