[TechKnowledge NOW] Grok 4.6 Released…Ties with GPT, 4.7 Also Announced

Photo of author

By Global Team

Elon Musk’s AI company SpaceXAI launched a new AI model, Grok 4.6, on the 12th local time. The model is designed to perform complex tasks that require multiple steps, such as research, information analysis, coding, and application development, from start to finish.

(Photo = Grok)
(Photo = Grok)

◆ From an AI that answers questions to one that takes on work

The market Grok 4.6 is targeting is agentic AI. Unlike conventional chatbots that provide a single answer to a single question, it refers to a method in which a model is given a goal, makes a plan, uses tools, checks intermediate results, and continues working without human intervention.

SpaceXAI said Grok 4.6 showed self-verifying behavior in long tasks. Before moving on to the next step, it tests and confirms what it has done. It also said the model is strong at tasks such as taking a product idea, researching an unfamiliar field, organizing the structure, implementing core features, and producing a working first version.

Key AI benchmark comparison results for Grok 4.6 High. (Source = SpaceXAI)
Key AI benchmark comparison results for Grok 4.6 High. (Source = SpaceXAI)

Its memory capacity has also expanded. Grok 4.6 can handle the equivalent of 500,000 tokens at once. A token is the smallest unit AI uses to process text, and 500,000 tokens is equivalent to several long novels. That means it can be fed a thick stack of contracts or an entire large software codebase and be assigned work on the whole thing.

The price has been kept the same as the previous version. Under the API, it costs $2 per 1 million input tokens and $6 per 1 million output tokens. The strategy appears to be a bet on price-performance competitiveness, with better performance at a frozen price. SpaceXAI is also offering double the usual usage amount in its own development tools for one week after launch as a promotion.

◆ The scorecard: No. 1 in documents, weaker in coding

AI model performance is compared through benchmark tests, standard exams that are similar to taking multiple subject tests and averaging the results.

In the Artificial Analysis Intelligence Index, the overall metric, Grok 4.6 scored 61 points. That was up 5 points from Grok 4.5 and matched OpenAI’s top-end GPT-5.6. According to VentureBeat, Anthropic remains at the top of the field. Claude Opus 5 is ranked No. 1, Claude Fabble 5 is second with 62 points, and Grok 4.6 is in a tie for third place with GPT-5.6.

The Grok 4.6 model scored 61 points, on par with GPT-5.6 Sol Max, and outperformed competing models in some evaluations. (Source = SpaceXAI)
The Grok 4.6 model scored 61 points, on par with GPT-5.6 Sol Max, and outperformed competing models in some evaluations. (Source = SpaceXAI)

The differences by category are large. In tests evaluating practical document writing and knowledge work, Grok 4.6 ranked first among the models compared. By contrast, in hands-on coding evaluations that fix actual software bugs, it scored 65.9 percent, behind GPT-5.6 (73 percent) and Claude Fabble 5 (70 percent). In terminal-based tasks, it was only about two-thirds as strong as competing models.

In short, it is strong for report writing and office work, but relatively weaker in demanding coding tasks. The result highlights an era in which different models are chosen depending on the use case.

◆ 4.7 teased for a few weeks later, with the upside and downside of speed

Musk has already revealed the schedule for the follow-up. According to multiple foreign media reports, including Loic News, Musk said Grok 4.7 would be released a few weeks after Grok 4.6. There are also predictions that Grok 4.7 will have 210 billion parameters. Parameters are the control units inside an AI “brain,” and in general, the more there are, the more complex the reasoning can be.

If the plan holds, it would mean a midsummer speed race in which two major models are launched one after the other. The expansion of the massive GPU cluster Colossus that SpaceXAI is building in Memphis, Tennessee, is being cited as the foundation for that speed.

There are also concerns about the fast release cycle. Foreign media reported warnings among AI safety researchers that release speed could outpace verification. Because Grok chatbots have previously drawn criticism over inappropriate responses, some analysts say that for corporate customers, trust management may matter as much as performance in securing adoption.

For users, intensifying model competition means more choices. As the gap between top-tier models narrows to just one or two points, price and use-case strengths are becoming key decision factors. If document work is the priority, choose Grok; if coding work is the priority, choose another model. Companies that delegate work to AI agents are also being advised to establish procedures for verifying the output.

If Grok 4.7 arrives as announced, the next round of summer AI competition will begin. The battleground is shifting beyond score comparisons to how long the models can work and how reliably they can do it.