What's new in Grok 4.6 from 500K context to pricing - #3169
Conversation
Appwrite WebsiteProject ID: Website (appwrite/website)Project ID: Tip Silent mode disables those chatty PR comments if you prefer peace and quiet |
Greptile SummaryAdds a new SEO blog post covering Grok 4.6 features, benchmarks, pricing, and Appwrite integration.
Confidence Score: 5/5The PR appears safe to merge. No blocking failure remains. Important Files Changed
Reviews (6): Last reviewed commit: "Revise pricing details in Grok 4.6 blog ..." | Re-trigger Greptile |
|
|
||
| **Training.** Grok 4.6 went through a longer supplemental training run than Grok 4.5, using curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe. SpaceXAI then used Grok 4.5 itself to regenerate the supervised fine-tuning trajectories across reasoning efforts, agent harnesses, and domains including STEM, software engineering, and knowledge work, filtering out problematic traces with model-based checks. Reinforcement learning ran on agentic tasks including general coding, knowledge work, kernel optimization, web development, and computer-aided design. | ||
|
|
||
| **Longer trajectories.** SpaceXAI says the model stays with complex tasks across many steps, and that on longer runs it started to see more self-testing and verification, with the model checking its own work before moving on. That behavior is the difference between an agent you can leave running and one you have to babysit. |
There was a problem hiding this comment.
| **Longer trajectories.** SpaceXAI says the model stays with complex tasks across many steps, and that on longer runs it started to see more self-testing and verification, with the model checking its own work before moving on. That behavior is the difference between an agent you can leave running and one you have to babysit. | |
| **Longer trajectories.** SpaceXAI says the model stays with complex tasks across many steps, with more self-testing and verification appearing on longer runs. |
| | Input modalities | Text and image | | ||
| | Output modalities | Text | | ||
| | Knowledge cutoff | February 1, 2026 | | ||
| | Reasoning | Configurable effort | |
There was a problem hiding this comment.
Don't think configurable efforts is the grounds to add it in table, it's better to just remove this one
|
|
||
| Here is the summary of what actually changed: | ||
|
|
||
| | Area | Grok 4.5 | Grok 4.6 | |
There was a problem hiding this comment.
I don't think the table is serving much purpose here, as we already cover Artificial analysis in the next heading, and primary focus is not really the major deal
| | Grok 4.6 | $2 | $6 | 500K | | ||
| | Grok 4.6 fast | $4 | $12 | 500K | | ||
| | GPT-5.6 Sol | $5 | $30 | 1M | | ||
| | Claude Opus 4.8 | $5 | $25 | 1M | |
There was a problem hiding this comment.
Opus 4.8 and Gemini 3.1 Pro are not a good comparison, these are outdated models, let's use Fable 5 and Opus 5 instead.
|
|
||
| The output side is where the gap shows. Agents spend most of their tokens writing: plans, tool arguments, diffs, test summaries, review notes. Moving from $25 or $30 per million output tokens to $6 changes which long tasks are worth automating at all, which is the same argument that made [Grok 4.5 interesting on cost](/blog/post/grok-45-coding-model). | ||
|
|
||
| # Independent measurements: what Grok 4.6 costs in latency |
There was a problem hiding this comment.
I'd suggest removing this section. Table already contains what we already had above. And Grok is probably the fastest model out of OpenAI, Anthropic and xAI.
Co-authored-by: Atharva Deosthale <atharva.deosthale17@gmail.com>
Co-authored-by: Atharva Deosthale <atharva.deosthale17@gmail.com>
Co-authored-by: Atharva Deosthale <atharva.deosthale17@gmail.com>
Updated pricing table to reflect new models and costs.


Latest SEO blog