Description
(I took the time to research/generate this for my own usage, sharing in case it's useful to someone else)
Model Recommendations for opencode Agent Roles - May 2026
Purpose
Recommend the top models for each opencode/oh-my-opencode-slim agent role, ranked by quality, cost, and suitability as of May 2026. Each tier provides top-5 closed and top-5 open-weight options.
Methodology
Sources:
- Artificial Analysis: throughput, pricing, quality index
- LMSYS Chatbot Arena: Elo ratings
- SWE-bench Verified: software engineering task completion
- Terminal-Bench: agentic terminal performance
- GPQA Diamond: hard science reasoning
- BenchLM, TokenMIX, OpenRouter rankings: cross-platform aggregation
- Aider Polyglot, LM Studio, OpenRouter community benchmarks: coding signals
Extra context
Tier-to-variant mapping: low = simple, medium = standard, high = complex, max = reasoning.
Reasoning Tier (max variant)
Used by: Oracle, Council summarizer, Council delta.
Criteria: deep reasoning, multi-file code generation, agentic planning, scientific reasoning. Latency is secondary to quality.
Closed Source
| Rank |
Model |
Key Strengths |
Approx Pricing (per M tokens) |
Notes |
| 1 |
Claude Opus 4.7 |
Best multi-file coding, Terminal-Bench #1, strongest agentic tool use |
$15/$75 |
Anthropic flagship for complex software engineering. 200K context. |
| 2 |
GPT-5.5 |
Highest AA Index (60), Terminal-Bench #2, fastest reasoning-class inference |
$10/$40 |
Best throughput of any frontier model. Excels at agentic chaining. |
| 3 |
Gemini 3.1 Pro |
GPQA 94.3%, best reasoning per dollar, 2M context |
$2/$10 |
Best value in reasoning tier. Unmatched context window. |
| 4 |
Claude Mythos |
Experimental Anthropic variant, reportedly SWE-bench 83.1% |
~$15/$75 |
Pre-release. For early adopters running Anthropic's latest reasoning architecture. |
| 5 |
GPT-5.5 Pro |
Extended reasoning variant of GPT-5.5, highest SWE-bench (84.2%) |
$15/$60 |
Best raw SWE-bench score. Use Oracle when correctness must not fail. |
Open Weight
| Rank |
Model |
Key Strengths |
Approx Pricing (per M tokens) |
Notes |
| 1 |
DeepSeek V4 Pro |
MIT license, 80.6% SWE-bench |
$0.55/$0.87 |
Open-weight champion. Self-hostable. 128K context. |
| 2 |
Kimi K2.6 |
79.4% SWE-bench, Chinese + English, 128K context |
$0.50/$0.80 |
Moonshot AI's latest. Strong multilingual code reasoning. |
| 3 |
Qwen 3.7 Max |
GPQA 92.1%, multilingual |
$0.60/$1.00 |
Best for multilingual reasoning. 131K context. |
| 4 |
GLM-5.1 |
Zhipu AI flagship, competitive with Kimi on code |
$0.45/$0.75 |
Solid all-rounder on Chinese cloud providers. |
| 5 |
Llama 4 Maverick |
405B-class MoE, uncensored variant |
$0.50/$0.80 |
Apache 2.0. Best for self-hosting at scale. |
Complex Tier (high variant)
Used by: Council gamma.
Criteria: strong coding and tool use, lower latency requirement than reasoning tier. Balance quality and speed.
Closed Source
| Rank |
Model |
Key Strengths |
Approx Pricing (per M tokens) |
Notes |
| 1 |
Claude Sonnet 4.6 |
Best quality-to-latency ratio in Claude lineup |
$3/$15 |
Nearly Opus-quality at 1/5 the cost. |
| 2 |
GPT-5.4 |
Fast, 128K context, strong Terminal-Bench scores |
$2.50/$10 |
Less capable than GPT-5.5 but significantly cheaper. |
| 3 |
Gemini 3 Flash |
97% of Gemini 3.1 Pro quality at 1/4 the cost |
$0.50/$3 |
Near reasoning-tier quality at standard-tier prices. |
| 4 |
Grok 4.1 |
xAI's latest, 128K context, strong math/code reasoning |
$3/$12 |
Less proven on SWE-bench. Competitive on math and science. |
| 5 |
GPT-5.4 Pro |
Higher-correctness variant of GPT-5.4 |
$5/$15 |
Middle ground between Sonnet 4.6 and Opus 4.7 pricing. |
Open Weight
| Rank |
Model |
Key Strengths |
Approx Pricing (per M tokens) |
Notes |
| 1 |
DeepSeek V4 Pro |
Cross-listed from reasoning tier |
$0.55/$0.87 |
Versatile enough for both tiers. Self-host alongside reasoning workload. |
| 2 |
Kimi K2.6 |
Cross-listed. Excellent code at open-weight pricing |
$0.50/$0.80 |
Good second source for council diversity. |
| 3 |
MiMo-V2.5-Pro |
01.AI's latest, 128K context |
$0.40/$0.70 |
Budget option for complex code generation. |
| 4 |
Mistral Large 3 |
European languages, good tool use |
$0.40/$0.70 |
Apache 2.0. Best for EU compliance. |
| 5 |
Qwen 3.5-397B |
Alibaba MoE model |
$0.35/$0.60 |
Slightly behind Kimi and DS V4 but ahead of older open models. |
Standard Tier (medium/low variant)
Used by: Librarian, Explorer, Observer, Designer, Scout.
Criteria: low latency, cost-sensitive, high throughput. These roles are I/O-bound, not reasoning-bound.
Closed Source
| Rank |
Model |
Key Strengths |
Approx Pricing (per M tokens) |
Notes |
| 1 |
Gemini 3 Flash |
0.5s TTFT, 97% quality retention from Pro, 1M context |
$0.50/$3 |
Best standard-tier option. Excellent for librarians and explorers. |
| 2 |
Claude Haiku 4.5 |
Fastest Claude variant, 200K context |
$0.25/$1.25 |
Best tool-use in standard tier. Slightly slower than Gemini Flash. |
| 3 |
GPT-5.4 Mini |
Fast, 128K context, good for structured data extraction |
$0.15/$0.60 |
Best for high-volume scraping. |
| 4 |
GPT-5.4 Nano |
Smallest GPT, extremely fast |
$0.05/$0.20 |
Lowest latency of any closed model. Ideal for rapid-fire explorer queries. |
| 5 |
Gemini 3.1 Flash-Lite |
Budget Flash variant |
$0.15/$0.60 |
Best value for Observer and basic designer tasks. |
Open Weight
| Rank |
Model |
Key Strengths |
Approx Pricing (per M tokens) |
Notes |
| 1 |
DeepSeek V4 Flash |
79% SWE-bench at 1/10 the cost, 128K context |
$0.14/$0.28 |
Budget champion. Self-hostable. Good enough for most standard tasks. |
| 2 |
MiniMax M2.7 |
Hugging Face leader in several categories |
$0.10/$0.20 |
Excellent for RAG-heavy librarian tasks. |
| 3 |
DeepSeek V3.2 |
Previous-gen, proven reliability |
$0.10/$0.20 |
Fallback. Mature ecosystem and hosting options. |
| 4 |
Mistral Small 4 |
Apache 2.0, fast code summarization |
$0.08/$0.16 |
Best for self-hosted setups on modest hardware (8B class). |
| 5 |
Qwen 3.5-72B |
Balanced quality and cost |
$0.12/$0.25 |
Good default for designer tasks. |
Simple Tier (low variant)
Used by: Fixer, Council alpha.
Criteria: maximum speed, minimum cost. Single-file edits that should be near-instant.
Closed Source
| Rank |
Model |
Key Strengths |
Approx Pricing (per M tokens) |
Notes |
| 1 |
Gemini 3 Flash |
Cross-listed from standard tier |
$0.50/$3 |
Fast, cheap, and good quality. |
| 2 |
GPT-5.4 Nano |
0.2s TTFT, $0.05/M in |
$0.05/$0.20 |
Best latency. Perfect for Fixer's iteration loop. |
| 3 |
Gemini 3.1 Flash-Lite |
Extremely cheap |
$0.15/$0.60 |
Runs on every council invocation. |
| 4 |
Mistral Small 3.2 |
Fastest Mistral variant |
$0.10/$0.20 |
European option. Fast inference. |
| 5 |
GPT-5 Nano |
Minimal GPT-5 model |
$0.03/$0.12 |
Cheapest closed model. Council alpha only (diversity of thought). |
Open Weight
| Rank |
Model |
Key Strengths |
Approx Pricing (per M tokens) |
Notes |
| 1 |
DeepSeek V4 Flash |
Cross-listed from standard tier |
$0.14/$0.28 |
Fast, good, cheap. |
| 2 |
Phi-4 Mini |
Microsoft's 3.8B |
$0.03/$0.06 |
Best model under 5B parameters. Self-host on CPU. |
| 3 |
Gemma 4 E4B |
Google's 4B MoE |
$0.03/$0.06 |
Strongest in its weight class. |
| 4 |
GPT-OSS-20B |
Community fine-tune, 20B, MIT license |
$0.05/$0.10 |
Distilled from GPT-4 lineage. |
| 5 |
Qwen 2.5 Coder 7B |
Proven code specialist |
$0.04/$0.08 |
Mature ecosystem. Well-tested for simple code changes. |
Specialist Picks
Observer (Vision/Multimodal)
- Closed 1: Gemini 3.1 Pro - 2M context, best multimodal understanding
- Closed 2: GPT-5.5 Vision - good vision, better code from screenshots
- Closed 3: Claude Opus 4.7 - best detailed visual reasoning
- Open 1: Qwen 3.5-Omni - strongest open-weight multimodal
- Open 2: Kimi-VL2.6 - good UI understanding
- Open 3: Llama 4 Maverick (multimodal variant) - Meta's best open vision model
Designer (UI-to-Code / Visual Design)
- Closed 1: Claude Opus 4.7 - best UI-to-code conversion
- Closed 2: GPT-5.5 - fastest UI iteration
- Closed 3: Gemini 3.1 Pro - design system consistency
- Open 1: UI2Code^N - specialized for visual design
- Open 2: Qwen 3.5-Omni - UI component creation
- Open 3: DeepSeek V4 Pro - generalist strong at UI code
Librarian (Search / RAG)
- Closed 1: Gemini 3.1 Pro - 2M context for massive codebase search
- Closed 2: Gemini 3 Flash - 97% quality at 1/4 cost
- Closed 3: Claude Haiku 4.5 - best tool use for search chains
- Open 1: DeepSeek V4 Flash - cheap search-heavy workloads
- Open 2: MiniMax M2.7 - best retrieval quality among open-weight models
- Open 3: Qwen 3.5-72B - balance of search comprehension and cost
Quick-Reference Table
| Tier |
Closed #1 |
Closed #2 |
Closed #3 |
Open #1 |
Open #2 |
Open #3 |
| Reasoning |
Claude Opus 4.7 |
GPT-5.5 |
Gemini 3.1 Pro |
DeepSeek V4 Pro |
Kimi K2.6 |
Qwen 3.7 Max |
| Complex |
Claude Sonnet 4.6 |
GPT-5.4 |
Gemini 3 Flash |
DeepSeek V4 Pro |
Kimi K2.6 |
MiMo-V2.5-Pro |
| Standard |
Gemini 3 Flash |
Claude Haiku 4.5 |
GPT-5.4 Mini |
DeepSeek V4 Flash |
MiniMax M2.7 |
DeepSeek V3.2 |
| Simple |
Gemini 3 Flash |
GPT-5.4 Nano |
Gemini 3.1 Flash-Lite |
DeepSeek V4 Flash |
Phi-4 Mini |
Gemma 4 E4B |
Caveats
- Pricing is approximate and provider-dependent. Self-hosted open-weight costs exclude hardware.
- Rankings reflect May 2026. The model landscape shifts monthly.
- Claude Mythos is speculative. Include only if on the cutting edge.
- Open-weight rankings assume self-hosting. Cloud-hosted pricing varies.
- For council diversity, avoid using the same model for multiple councillors.
- Specialized models (Phi-4 Mini, Qwen 2.5 Coder 7B) may outperform generalists on their specific task.
Description
(I took the time to research/generate this for my own usage, sharing in case it's useful to someone else)
Model Recommendations for opencode Agent Roles - May 2026
Purpose
Recommend the top models for each opencode/oh-my-opencode-slim agent role, ranked by quality, cost, and suitability as of May 2026. Each tier provides top-5 closed and top-5 open-weight options.
Methodology
Sources:
Extra context
Tier-to-variant mapping:
low= simple,medium= standard,high= complex,max= reasoning.Reasoning Tier (max variant)
Used by: Oracle, Council summarizer, Council delta.
Criteria: deep reasoning, multi-file code generation, agentic planning, scientific reasoning. Latency is secondary to quality.
Closed Source
Open Weight
Complex Tier (high variant)
Used by: Council gamma.
Criteria: strong coding and tool use, lower latency requirement than reasoning tier. Balance quality and speed.
Closed Source
Open Weight
Standard Tier (medium/low variant)
Used by: Librarian, Explorer, Observer, Designer, Scout.
Criteria: low latency, cost-sensitive, high throughput. These roles are I/O-bound, not reasoning-bound.
Closed Source
Open Weight
Simple Tier (low variant)
Used by: Fixer, Council alpha.
Criteria: maximum speed, minimum cost. Single-file edits that should be near-instant.
Closed Source
Open Weight
Specialist Picks
Observer (Vision/Multimodal)
Designer (UI-to-Code / Visual Design)
Librarian (Search / RAG)
Quick-Reference Table
Caveats