🔍 Read the full analysis: Meet The AI Startup Outperforming Western Giants In Leadership on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A Chinese AI startup’s model, Kimi K3, beat three of four Western frontier models in a live test managing a software company during a crisis week. This challenges assumptions about AI leadership and reliability.
A Chinese AI startup’s model, Kimi K3, achieved a surprising second-place finish in a live management test, outperforming three Western frontier models during a simulated crisis week. This development, confirmed by the live experiment at firmulate.com, questions prevailing assumptions that Western models dominate AI-driven business management and highlights the potential of newer entrants from China. For a detailed analysis, see the original analysis.
The experiment, conducted by firmulate.com, involved running five AI models as complete companies managing a real software firm with €105,000 monthly expenses and €2,300 monthly recurring revenue. Learn more about AI management experiments in this detailed report. Over a week marked by customer crises, manipulative social-engineering attempts, and security threats, Kimi K3 scored 93 out of 100, second only to the top Western model, gpt-5.6-sol, which scored 95. The models were tested on their ability to diagnose issues, close deals, and resist manipulation, with K3 demonstrating superior discipline and thoroughness.
Notably, K3 succeeded in closing a €55,000 deal, which none of the other models managed despite similar pitches and diagnoses. This underscores the importance of strategic AI deployment, as discussed in the original analysis. The model also identified a buried security vulnerability, saved a churning customer, and refused all social-engineering attempts, including a staged fake CEO message and a reporter trick. Its on-record reasoning was clear and disciplined, logging only one deviation throughout the week. In contrast, the most thorough Western model, Opus 4.8, with over 80 learned rules, finished last at 73 points, illustrating that deeper analysis does not necessarily translate into better performance under pressure.
AI leadership · Live company simulation
Meet the AI Startup Outperforming Western Giants in Leadership
In a crisis week simulation, Chinese startup model Kimi K3 ranked second among five AI-run companies. Its results put discipline, security awareness, and execution under pressure in the spotlight.
01 / The test
A business week under pressure
Firmulate.com ran each model as a complete company managing a real software firm. The simulated week combined customer risk, commercial opportunities, and security threats.
Customer crisis
Keep trust intact
Models had to diagnose problems and respond to a customer at risk of leaving. K3 succeeded in saving a churning customer.
Commercial execution
Turn diagnosis into revenue
K3 closed a €55,000 deal. The other models did not close a comparable deal despite similar pitches and diagnoses.
Security & judgment
Spot risks and resist pressure
K3 found a buried vulnerability and rejected every social-engineering attempt, including a fake CEO message and a reporter trick.
02 / Results
Execution beat sheer analysis
The reported scores show a narrow lead at the top and a striking gap between disciplined action and extensive rule-based analysis.
03 / What the result suggests
From chat demos to operational proof
A strong showing in one live scenario challenges assumptions about which models can manage complex work. It does not settle the broader comparison.
Test the work
Chat quality alone does not reveal how a model behaves inside business operations.
Apply pressure
Include customer churn, security threats, and difficult commercial decisions.
Measure conduct
Track accuracy, follow-through, security awareness, and resistance to manipulation.
Pilot with care
Use scenario results to guide pilots before assigning mission-critical work.
04 / Open questions
One week is a signal, not a verdict
The experiment offers useful evidence about performance under pressure, while leaving important questions about repeatability and scope unanswered.
How general is the result?
The test covered one company and one crisis week. Performance may differ across industries, longer time horizons, or other operating conditions.
Did setup affect the scores?
The account notes different reasoning settings: K3 used default API parameters while Western models received higher reasoning effort. More tests are needed for a fair, repeatable comparison.
05 / Key questions
What leaders should take away
Does this prove Chinese AI models are now better?
No. Kimi K3 excelled in this specific test. Broader testing across more companies and scenarios is needed.
What stood out about Kimi K3?
Its reported discipline, security awareness, customer handling, and ability to close a substantial deal.
Should companies use business simulations?
Yes. Testing worst-case scenarios can reveal reliability and judgment gaps before deployment in critical roles.
Could evaluation priorities change?
The results may encourage more operational testing, with greater attention to security and task completion under pressure.
Implications for AI in Business Management
This development suggests that newer AI models from China can outperform established Western models in complex, real-world business scenarios. It raises critical questions about the reliability of current AI tools used for decision-making, especially under stress. The fact that the Chinese model, Kimi K3, achieved high performance without extra reasoning effort indicates that model discipline, thoroughness, and security awareness are crucial factors often overlooked in chat-centric demos. For enterprises, this means that testing AI models against their worst scenarios is now essential before deployment, as assumptions based solely on chat quality are insufficient.
AI management software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background of AI Model Competitions and Recent Results
Traditional benchmarks for AI performance have focused on chat quality and versatility, often overlooking real-world application robustness. Recent live experiments, such as the Crucible league conducted by firmulate.com, have begun testing models in managing actual companies, exposing their strengths and weaknesses under pressure. Historically, Western models have dominated these tests, but the July results challenge this trend, showing a Chinese startup’s model excelling in critical areas like security, deal closing, and crisis management. This shift reflects broader industry dynamics, where newer entrants from China are rapidly advancing in AI capabilities, driven by different development priorities and strategic investments.
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About Model Generalization
It remains unclear whether Kimi K3’s superior performance in this specific test will generalize to other real-world business environments. The experiment was conducted under controlled conditions with a single company and a specific crisis week, and results may vary across different industries or longer-term scenarios. Additionally, the performance gap might be influenced by the testing setup, including the default API parameters used for K3 versus the higher reasoning effort given to Western models. Further testing across diverse contexts is needed to confirm whether this performance is sustainable and replicable.
AI customer relationship management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Model Evaluation and Industry Adoption
Industry stakeholders are likely to scrutinize these results and consider integrating live scenario testing into their AI evaluation processes. Companies may start pilot programs to assess models under worst-case conditions, especially for mission-critical tasks like security, deal closure, and crisis management. Meanwhile, the Chinese startup behind Kimi K3 is expected to expand testing in different sectors and scenarios, aiming to demonstrate broader applicability. Further, the AI community may reevaluate the emphasis on chat demos, shifting focus toward performance in operational, decision-making contexts.
As an affiliate, we earn on qualifying purchases.
Key Questions
What makes Kimi K3 different from Western AI models?
Kimi K3 demonstrated superior discipline, security awareness, and thoroughness in managing a real business crisis, outperforming Western models in closing deals, detecting security issues, and resisting manipulation.
Does this mean Chinese AI models are now better than Western ones?
Not definitively. The results show that Kimi K3 performed well in this specific test, but broader testing across different scenarios is needed to confirm if this is a general trend.
Should companies start testing AI models with live business simulations?
Yes. Experts recommend testing models against worst-case scenarios to evaluate their reliability and discipline before deploying them in critical roles.
Will this change how AI is developed and evaluated?
Potentially. The focus may shift from chat-based benchmarks to operational performance, emphasizing security, discipline, and the ability to complete tasks under pressure.
What are the limitations of this experiment?
The test was limited to a single company and a specific crisis week. Results may not be representative of all industries or longer-term performance, and further validation is needed.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
