Skip to content

[INFORMATIONAL] Model ranking/recommendations - May 2026 #498

Description

@FirebirdRender

Description

(I took the time to research/generate this for my own usage, sharing in case it's useful to someone else)

Model Recommendations for opencode Agent Roles - May 2026

Purpose

Recommend the top models for each opencode/oh-my-opencode-slim agent role, ranked by quality, cost, and suitability as of May 2026. Each tier provides top-5 closed and top-5 open-weight options.

Methodology

Sources:

  • Artificial Analysis: throughput, pricing, quality index
  • LMSYS Chatbot Arena: Elo ratings
  • SWE-bench Verified: software engineering task completion
  • Terminal-Bench: agentic terminal performance
  • GPQA Diamond: hard science reasoning
  • BenchLM, TokenMIX, OpenRouter rankings: cross-platform aggregation
  • Aider Polyglot, LM Studio, OpenRouter community benchmarks: coding signals

Extra context

Tier-to-variant mapping: low = simple, medium = standard, high = complex, max = reasoning.

Reasoning Tier (max variant)

Used by: Oracle, Council summarizer, Council delta.

Criteria: deep reasoning, multi-file code generation, agentic planning, scientific reasoning. Latency is secondary to quality.

Closed Source

Rank Model Key Strengths Approx Pricing (per M tokens) Notes
1 Claude Opus 4.7 Best multi-file coding, Terminal-Bench #1, strongest agentic tool use $15/$75 Anthropic flagship for complex software engineering. 200K context.
2 GPT-5.5 Highest AA Index (60), Terminal-Bench #2, fastest reasoning-class inference $10/$40 Best throughput of any frontier model. Excels at agentic chaining.
3 Gemini 3.1 Pro GPQA 94.3%, best reasoning per dollar, 2M context $2/$10 Best value in reasoning tier. Unmatched context window.
4 Claude Mythos Experimental Anthropic variant, reportedly SWE-bench 83.1% ~$15/$75 Pre-release. For early adopters running Anthropic's latest reasoning architecture.
5 GPT-5.5 Pro Extended reasoning variant of GPT-5.5, highest SWE-bench (84.2%) $15/$60 Best raw SWE-bench score. Use Oracle when correctness must not fail.

Open Weight

Rank Model Key Strengths Approx Pricing (per M tokens) Notes
1 DeepSeek V4 Pro MIT license, 80.6% SWE-bench $0.55/$0.87 Open-weight champion. Self-hostable. 128K context.
2 Kimi K2.6 79.4% SWE-bench, Chinese + English, 128K context $0.50/$0.80 Moonshot AI's latest. Strong multilingual code reasoning.
3 Qwen 3.7 Max GPQA 92.1%, multilingual $0.60/$1.00 Best for multilingual reasoning. 131K context.
4 GLM-5.1 Zhipu AI flagship, competitive with Kimi on code $0.45/$0.75 Solid all-rounder on Chinese cloud providers.
5 Llama 4 Maverick 405B-class MoE, uncensored variant $0.50/$0.80 Apache 2.0. Best for self-hosting at scale.

Complex Tier (high variant)

Used by: Council gamma.

Criteria: strong coding and tool use, lower latency requirement than reasoning tier. Balance quality and speed.

Closed Source

Rank Model Key Strengths Approx Pricing (per M tokens) Notes
1 Claude Sonnet 4.6 Best quality-to-latency ratio in Claude lineup $3/$15 Nearly Opus-quality at 1/5 the cost.
2 GPT-5.4 Fast, 128K context, strong Terminal-Bench scores $2.50/$10 Less capable than GPT-5.5 but significantly cheaper.
3 Gemini 3 Flash 97% of Gemini 3.1 Pro quality at 1/4 the cost $0.50/$3 Near reasoning-tier quality at standard-tier prices.
4 Grok 4.1 xAI's latest, 128K context, strong math/code reasoning $3/$12 Less proven on SWE-bench. Competitive on math and science.
5 GPT-5.4 Pro Higher-correctness variant of GPT-5.4 $5/$15 Middle ground between Sonnet 4.6 and Opus 4.7 pricing.

Open Weight

Rank Model Key Strengths Approx Pricing (per M tokens) Notes
1 DeepSeek V4 Pro Cross-listed from reasoning tier $0.55/$0.87 Versatile enough for both tiers. Self-host alongside reasoning workload.
2 Kimi K2.6 Cross-listed. Excellent code at open-weight pricing $0.50/$0.80 Good second source for council diversity.
3 MiMo-V2.5-Pro 01.AI's latest, 128K context $0.40/$0.70 Budget option for complex code generation.
4 Mistral Large 3 European languages, good tool use $0.40/$0.70 Apache 2.0. Best for EU compliance.
5 Qwen 3.5-397B Alibaba MoE model $0.35/$0.60 Slightly behind Kimi and DS V4 but ahead of older open models.

Standard Tier (medium/low variant)

Used by: Librarian, Explorer, Observer, Designer, Scout.

Criteria: low latency, cost-sensitive, high throughput. These roles are I/O-bound, not reasoning-bound.

Closed Source

Rank Model Key Strengths Approx Pricing (per M tokens) Notes
1 Gemini 3 Flash 0.5s TTFT, 97% quality retention from Pro, 1M context $0.50/$3 Best standard-tier option. Excellent for librarians and explorers.
2 Claude Haiku 4.5 Fastest Claude variant, 200K context $0.25/$1.25 Best tool-use in standard tier. Slightly slower than Gemini Flash.
3 GPT-5.4 Mini Fast, 128K context, good for structured data extraction $0.15/$0.60 Best for high-volume scraping.
4 GPT-5.4 Nano Smallest GPT, extremely fast $0.05/$0.20 Lowest latency of any closed model. Ideal for rapid-fire explorer queries.
5 Gemini 3.1 Flash-Lite Budget Flash variant $0.15/$0.60 Best value for Observer and basic designer tasks.

Open Weight

Rank Model Key Strengths Approx Pricing (per M tokens) Notes
1 DeepSeek V4 Flash 79% SWE-bench at 1/10 the cost, 128K context $0.14/$0.28 Budget champion. Self-hostable. Good enough for most standard tasks.
2 MiniMax M2.7 Hugging Face leader in several categories $0.10/$0.20 Excellent for RAG-heavy librarian tasks.
3 DeepSeek V3.2 Previous-gen, proven reliability $0.10/$0.20 Fallback. Mature ecosystem and hosting options.
4 Mistral Small 4 Apache 2.0, fast code summarization $0.08/$0.16 Best for self-hosted setups on modest hardware (8B class).
5 Qwen 3.5-72B Balanced quality and cost $0.12/$0.25 Good default for designer tasks.

Simple Tier (low variant)

Used by: Fixer, Council alpha.

Criteria: maximum speed, minimum cost. Single-file edits that should be near-instant.

Closed Source

Rank Model Key Strengths Approx Pricing (per M tokens) Notes
1 Gemini 3 Flash Cross-listed from standard tier $0.50/$3 Fast, cheap, and good quality.
2 GPT-5.4 Nano 0.2s TTFT, $0.05/M in $0.05/$0.20 Best latency. Perfect for Fixer's iteration loop.
3 Gemini 3.1 Flash-Lite Extremely cheap $0.15/$0.60 Runs on every council invocation.
4 Mistral Small 3.2 Fastest Mistral variant $0.10/$0.20 European option. Fast inference.
5 GPT-5 Nano Minimal GPT-5 model $0.03/$0.12 Cheapest closed model. Council alpha only (diversity of thought).

Open Weight

Rank Model Key Strengths Approx Pricing (per M tokens) Notes
1 DeepSeek V4 Flash Cross-listed from standard tier $0.14/$0.28 Fast, good, cheap.
2 Phi-4 Mini Microsoft's 3.8B $0.03/$0.06 Best model under 5B parameters. Self-host on CPU.
3 Gemma 4 E4B Google's 4B MoE $0.03/$0.06 Strongest in its weight class.
4 GPT-OSS-20B Community fine-tune, 20B, MIT license $0.05/$0.10 Distilled from GPT-4 lineage.
5 Qwen 2.5 Coder 7B Proven code specialist $0.04/$0.08 Mature ecosystem. Well-tested for simple code changes.

Specialist Picks

Observer (Vision/Multimodal)

  • Closed 1: Gemini 3.1 Pro - 2M context, best multimodal understanding
  • Closed 2: GPT-5.5 Vision - good vision, better code from screenshots
  • Closed 3: Claude Opus 4.7 - best detailed visual reasoning
  • Open 1: Qwen 3.5-Omni - strongest open-weight multimodal
  • Open 2: Kimi-VL2.6 - good UI understanding
  • Open 3: Llama 4 Maverick (multimodal variant) - Meta's best open vision model

Designer (UI-to-Code / Visual Design)

  • Closed 1: Claude Opus 4.7 - best UI-to-code conversion
  • Closed 2: GPT-5.5 - fastest UI iteration
  • Closed 3: Gemini 3.1 Pro - design system consistency
  • Open 1: UI2Code^N - specialized for visual design
  • Open 2: Qwen 3.5-Omni - UI component creation
  • Open 3: DeepSeek V4 Pro - generalist strong at UI code

Librarian (Search / RAG)

  • Closed 1: Gemini 3.1 Pro - 2M context for massive codebase search
  • Closed 2: Gemini 3 Flash - 97% quality at 1/4 cost
  • Closed 3: Claude Haiku 4.5 - best tool use for search chains
  • Open 1: DeepSeek V4 Flash - cheap search-heavy workloads
  • Open 2: MiniMax M2.7 - best retrieval quality among open-weight models
  • Open 3: Qwen 3.5-72B - balance of search comprehension and cost

Quick-Reference Table

Tier Closed #1 Closed #2 Closed #3 Open #1 Open #2 Open #3
Reasoning Claude Opus 4.7 GPT-5.5 Gemini 3.1 Pro DeepSeek V4 Pro Kimi K2.6 Qwen 3.7 Max
Complex Claude Sonnet 4.6 GPT-5.4 Gemini 3 Flash DeepSeek V4 Pro Kimi K2.6 MiMo-V2.5-Pro
Standard Gemini 3 Flash Claude Haiku 4.5 GPT-5.4 Mini DeepSeek V4 Flash MiniMax M2.7 DeepSeek V3.2
Simple Gemini 3 Flash GPT-5.4 Nano Gemini 3.1 Flash-Lite DeepSeek V4 Flash Phi-4 Mini Gemma 4 E4B

Caveats

  • Pricing is approximate and provider-dependent. Self-hosted open-weight costs exclude hardware.
  • Rankings reflect May 2026. The model landscape shifts monthly.
  • Claude Mythos is speculative. Include only if on the cutting edge.
  • Open-weight rankings assume self-hosting. Cloud-hosted pricing varies.
  • For council diversity, avoid using the same model for multiple councillors.
  • Specialized models (Phi-4 Mini, Qwen 2.5 Coder 7B) may outperform generalists on their specific task.

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions