Researchers pit ChatGPT, Claude, Gemini, Grok, and DeepSeek in a Pokemon tournament

Researchers organized a competitive Pokemon tournament pitting five major large language models against one another: ChatGPT, Claude, Gemini, Grok, and DeepSeek. The benchmark tests whether frontier AI systems can integrate multiple simultaneous constraints during strategic play. Pokemon battles demand simultaneous evaluation of type matchups, stat trade-offs, move power, and move accuracy, making the game a multi-variable reasoning challenge rather than a single-turn puzzle. The public report, shared on X by @thesupermanmx, describes the setup and the cognitive demands of the task but does not yet disclose match outcomes, scoring methodology, or independent verification. For builders evaluating model selection, the tournament framing signals growing interest in game-based evals that stress combinatorial planning under uncertainty. Readers should treat this as an early signal until fuller results or peer-reviewed reporting emerges.
Sources
- Researchers pit ChatGPT, Claude, Gemini, Grok, and DeepSeek in a Pokemon tournament@thesupermanmx on X · primary