Introducing Grok 4.7

xAI released Grok 4.7, a larger base model tuned for longer coding tasks and self‑verification, priced the same as Grok 4.6. The announcement cites superior CursorBench 4.0 price‑performance and higher scores on DeepSWE, EEBench and legal benchmarks, while adding an upgraded safeguard stack. No public weights are linked yet.

xAI announced the release of Grok 4.7, a new language model designed for extended coding sessions and professional knowledge tasks. The company stated the system uses a larger base architecture than its predecessor, Grok 4.6, and underwent a longer reinforcement learning phase. This training emphasized difficult problems that require many hours to solve, aiming to improve self-verification skills. The model is currently available through various APIs and coding environments at the same price as the previous version. Critics often judge new AI systems by their ability to handle complex, long-running workflows without losing context or accuracy. xAI claims Grok 4.7 achieves frontier-level price-performance on the CursorBench 4.0 test, which specifically stresses these longer coding tasks. The provider also highlighted improved scores on DeepSWE, EEBench, and legal benchmarks compared to the older model. Additionally, the new safeguard stack reportedly offers stronger resistance to jailbreaks while maintaining utility for benign cybersecurity research. The announcement does not link to public model weights, so independent verification of the underlying architecture remains impossible at this time. Several benchmark results, particularly for DeepSWE, are marked as high-effort scores, which may not reflect standard operational consistency. While the publisher describes the safeguards as the strongest tested, third-party audits of the new cyber defense capabilities have not yet been published. Comparisons to rival systems like GPT-5.6 Sol and Fable 5.1 are based solely on provider-selected metrics.