
VigilSAR, a specialized defense-ISR software platform, has released its latest public LLM leaderboard—a crucial tool for assessing language models in intelligence, surveillance, and reconnaissance tasks. Unlike typical benchmarks, this one emphasizes reasoning, reporting, and restraint over general trivia, making it highly relevant for defense applications.
The evaluation involved 14 models tested across 300 tasks, with results recorded on July 17, 2026. Importantly, the task set is private, deliberately kept secret to prevent models from training on it and to ensure genuine performance assessment. A separate, private held-out set exists, with the gap between public and held-out scores published per model, serving as an indicator of potential memorization or overfitting.
Current standings are organized by confidence bands rather than precise ranks, emphasizing the reliability of each model’s performance. Leading the pack is Claude-Fable-5 with a score of 67.77 in Band A. A noteworthy new entry is Moonshot’s Kimi K3, debuting at #3 with a score of 64.65 in Band B. Kimi K3 surpasses all GPT and Gemini models on the leaderboard, marking a significant breakthrough for locally-runnable, deployable models in defense contexts.
The scoring framework also incorporates cost-per-correct-answer economics and deployment considerations. The evaluation aims to determine which models are trusted enough for real-world defense tasks, with no vendor claims influencing the process—only measurable performance. This approach fosters transparency and honesty in a space often clouded by vendor marketing claims.
At its core, the site emphasizes the importance of bands over ranks, complemented by confidence intervals and the published public leaderboard. The system’s design ensures that results are meaningful and robust, especially given the private task set that preserves the integrity of the evaluation against training contamination.

For tech enthusiasts and defense analysts alike, VigilSAR’s benchmark represents a significant step toward trustworthy evaluation of LLMs in sensitive applications. The debut of Kimi K3 at #3, ahead of prominent GPT and Gemini models, highlights the rapid evolution of locally-runnable, deployable models capable of meeting complex ISR demands. To see the latest standings in detail, visit the public leaderboard or learn more about the platform at VigilSAR.
defense LLM benchmarking software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
ISR AI model deployment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
locally runnable AI models for defense
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI reasoning and reporting tools for defense
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.