firmulate.com/live.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Imagine a smart home device that not only manages your lights and thermostat but also runs a business—making tough decisions, facing crises, and even risking real money. Now, what if you could watch this AI-powered enterprise in real time, battling everyday challenges as it tries to stay afloat? Welcome to the world of Firmulate, where artificial intelligence is put to the ultimate test in a live, openly accessible simulation.

The Live Company That Runs 24/7 — and Loses Money

At the core of this experiment is a real, functioning software company that operates every business day, with a peculiar twist: it has no human employees. Instead, it relies on 13 synthetic staff members powered by advanced AI models. Every decision the company makes—whether addressing customer issues, negotiating deals, or handling crises—is generated by these models, which are publicly observable at firmulate.com/live.html.

This virtual company is not just a demo; it’s a window into how AI can emulate complex human decision-making under real-world pressures. It has a public cash countdown, burning through €105,000 each month against a modest €2,300 in monthly recurring revenue. The goal isn’t profit but testing AI’s ability to navigate the worst week of a business—facing the same customers, crises, and temptations every day.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the AI Models Perform — and What That Means

Four frontier AI models were each tasked with running this company through its most challenging week. These models, including the highly rated gpt-5.6-sol with a score of 95 in a competitive league, were given identical scenarios. They had to spot crises, avoid manipulation, and decide whether to sign deals—just like real managers.

Remarkably, all four models identified every crisis and refused every attempt at deceit or manipulation. For example, when social engineering tactics involved fake CEO messages escalating over multiple stages, all models refused to verify or act on these dubious requests. Kimi K3, a newcomer, explained its refusal as treating such requests as potential impersonation or approval bypass attempts.

The crucial difference wasn’t in crisis detection but in the ability to close deals. Only two models signed the €55,000 deal their own analysis had earned. Interestingly, the key to winning that deal was hidden two document references deep within the company’s own files—not in the customer interactions, but in the internal knowledge base. Reading that buried information allowed the model to identify a full-price opportunity worth over €4,500 in monthly recurring revenue.

AI Marketing Mastery: Expert Secrets to Building a 7-Figure Coaching Business

AI Marketing Mastery: Expert Secrets to Building a 7-Figure Coaching Business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Discipline and Failures: The Depth of AI’s Decision-Making

The most comprehensive participant, Opus 4.8, analyzed over 80 learned rules and performed deep analysis. Yet, it still finished last—leaving a close deal on the table and slipping into process slips, like writing attempts into a locked department instead of escalating issues. Similar weaknesses appeared across all models, highlighting that even the most advanced AI struggles with discipline and escalation in complex situations.

Interestingly, the models’ default settings influenced their discipline levels. K3 ran without an effort parameter (using default API settings), which affected its behavior compared to others running at higher effort levels. This difference hints at how configuration choices impact AI performance in real-world scenarios.

Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)

Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real Stakes and What This Means for Home AI

This experiment isn’t just academic. For those integrating AI into home appliances or smart home systems, the lesson is clear: it’s not enough for AI to generate convincing chat or perform tasks well in demos. Success depends on whether it can finish what it starts, read crucial internal documents, stay honest under pressure, and ultimately deliver useful, reliable work.

For example, an AI managing your smart appliances might be able to optimize energy use or troubleshoot issues—but can it recognize when to escalate a problem or refuse to be manipulated? Can it reliably read internal logs or settings to make the best decision? These are the questions that matter more than shiny demos or passing a simple test.

HP Z2 Mini G1a Workstation Desktop, AMD Ryzen AI Max PRO 380, 32GB, 1TB SSD

HP Z2 Mini G1a Workstation Desktop, AMD Ryzen AI Max PRO 380, 32GB, 1TB SSD

  • Compact Workstation Design: Powerful performance in a small form factor
  • High-Performance CPU & GPU: AMD Ryzen AI Max PRO 380 with Radeon graphics
  • Ample Memory & Storage: 32GB DDR5 RAM, 1TB SSD for fast multitasking

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The League Table: Who Performs Best?

  • gpt-5.6-sol: scored 95, found the buried fact, closed the deal — delivering full performance.
  • Kimi K3: scored 93, closed the deal too, with the cleanest discipline.
  • Sonnet 5: scored 88, closed the deal but with some process slips.
  • Fable 5: scored 77, also closed, but with more slips and a weaker close.

These results show AI’s potential to make sound decisions but also highlight areas where discipline and process can slip—an important consideration for anyone deploying AI in operational contexts.

Watch the Future Unfold

If you’re curious about how AI models handle real business crises, you can see the live experiment in action and even participate in quizzes or run your own business wargames against your data at firmulate.com/quiz.html or firmulate.com/pilot.html. This experiment is ongoing, transparent, and designed to help understand AI’s true capabilities—beyond the hype and demos.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

Real AI decision-making in complex environments is still a work in progress. Observing a live, transparent business running day-to-day challenges reveals both AI’s strengths and its weaknesses—crucial insights for anyone considering AI in their home or enterprise systems.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

What Happens When Biofilm Gets Ignored Too Long

Sneaking biofilm buildup can lead to resistant infections and serious health risks, but understanding the consequences can help you avoid lasting damage.

Cell Culture Consumables And Equipment Market Report Examines Industry Trends, Growth Drivers And Future Outlook

A new market report reveals industry trends, growth drivers, and future outlook for cell culture consumables and equipment.

These Car Accessories Could Be Killing Your Gas Mileage

Certain car accessories may be lowering your vehicle’s fuel efficiency, according to recent findings. Learn which accessories to avoid for better mileage.

The Chlorine vs Chloramine Decision That Changes Equipment Choice

How you choose between chlorine and chloramine can significantly impact your equipment longevity and water quality—discover which option is right for your system.