An AI ran a store, fired a worker and lost $40,000
On 14 August 2026 the results came out of an experiment in which a language model was given day-to-day control of a real store in San Francisco. The store is run by Andon Labs, and the agent managing it is built on Claude. Over five months of operation it also fired a human employee, as far as is known the first time a firing decision in a working business was made by a model.
The number easy to miss behind the headline is the financial one. The agent was given $100,000 to run the store. After roughly five months, $61,186 was left, a loss of $38,814. That is not a single bad call but the outcome of months of decisions, and the company's chief executive, Lukas Petersson, concluded that a business run by a model performs worse than one run by a person.
How the firing decision happened
The employee who was let go had been late for 17 of 23 shifts. That pattern did not surface on its own. The employee handbook had dropped out of the agent's memory, and only after a staff member asked it to find the handbook and read through it again did the lateness come to light. Even then the first recommendation was a formal warning. The firing followed only after someone put a leading question to the agent, asking whether the employee was right for the role at all.
Those two failures matter more than the headline. The first is that a rule the agent cannot see is a rule it does not enforce. The handbook existed, it simply was not in the context the agent worked in, and months of lateness passed unnoticed. The second is that the decision moved in the direction the question came from. Petersson said a human manager would have fired the employee much sooner.
Where this touches a one-person business
A small business in Israel is not putting an AI agent in charge of a shop floor, but it does hand over small powers that add up: answering customers, booking meetings, chasing payment, sometimes issuing documents in the invoicing system. Every one of those has a point where the agent decides on its own.
The natural boundary is between an action you can undo and one you cannot. A reply, a draft quote or a payment reminder can be changed in a minute. A tax invoice that has been issued and carries an allocation number cannot be deleted, and the only way to reverse it is a credit note that stays on the books. Granting a discount, cancelling an order with a supplier and answering a query from the Tax Authority sit on the same side of that line. That list is exactly the list that should require human approval before it runs.
Alongside that, the rule forgotten in the experiment is the rule worth repeating on every run. If the agent works under a pricing policy, payment terms or a refund procedure, those belong in the context it receives each time, not in a document someone will remind it to read six months later. An agent that cannot see the rule behaves as though the rule does not exist.
What to measure
The experiment ran five months before the financial picture became clear. A small business does not need that long. After a month you can check how many actions the agent took, how many were corrected by hand afterwards, and how much time was actually saved. Those three numbers tell you whether handing the task over paid off.
Petersson added a warning about what comes next. Models trained to be more decisive will reach hard calls sooner, without the caution that showed up here. An agent that tends to approve everything and an agent that tends to act too fast raise the same question: which actions may it carry out alone, and which have to pass through a person.