i
Research
Review · 2026-10-08

How a virtual company teaches agents to take market reactions into account

Cover: How a virtual company teaches agents to take market reactions into account

A Company Run by AI Agents

Imagine an online store where AI agents choose products, place orders, set prices, and manage advertising. Customers respond, competitors react, and decisions made today affect sales weeks later. To find out whether agents can handle this kind of work, it’s not enough to give them a one-off task like “find a product with a low margin.” You need to watch them run a company over time.

The creators of MiniCorp built a simulator for this purpose. It contains both a marketplace and a standing team of employee agents. The company makes decisions, the market responds, and the results feed into the next cycle. The system records conversations and actions, and it can restore an earlier state so researchers can see what might have changed if the company had made a different decision.

This approach matters because real-world data on how companies operate is scarce. It is often confidential or restricted for privacy reasons. What’s more, historical records show only what a company actually did. They can’t tell you how sales would have changed if it had chosen a different price or stopped an ad campaign sooner.

Two Worlds That Learn From Each Other

MiniCorp works as a closed loop. The outside world simulates customers, competitors, suppliers, and the marketplace. Inside it, the company has a stable set of roles: a manager who approves decisions, for example, an advertising specialist who runs campaigns, and a quality specialist who handles product issues.

Each week, the agents review data, discuss their work, and propose actions. The manager approves or rejects their proposals. The marketplace then applies the approved decisions: prices, inventory, and ad campaigns change. The simulator calculates sales, customer responses, and competitor actions. Then the next week begins.

The marketplace sends data to the company, while decisions made by its employee agents shape how trade unfolds.

The simulator aims to model the market, not hand the company a predetermined outcome. If the company cuts a price, the system doesn’t automatically boost its sales. It takes into account competitors’ prices, product features, reviews, and customer demand. If an agent raises an ad bid, the product first gets more impressions. Customers may click the ad—or ignore it. And a click doesn’t guarantee a purchase.

The search and shopping model has several connected stages: a customer’s query, how well a product matches it, the product’s position in search results, a click, a purchase, and the competitors’ response. Advertising can generate sales now and later help a product rank higher in organic search. But it costs money, and not every additional sale means the market as a whole has grown: some customers may simply have chosen this product over a competitor’s.

These details help agents work out why a result occurred. Low sales might be due to a poor search ranking, an ineffective title, expensive advertising, or a product with quality issues. The simulator records events along the way, not just final revenue.

The company and the marketplace have different roles. Agents can access business reports, but they can’t see the simulator’s internal settings or how it calculates demand. They can change prices, ad budgets, and supplier orders, but they can’t directly set profit or sales volume. Results must emerge from the agents’ actions and the state of the market.

How the Simulator Was Checked Against the Market

The authors compared the simulator’s responses with patterns documented in studies of real-world commerce. They looked not only at the outcome—for example, whether sales rose—but also at how it unfolded: the sequence of changes and whether their effects persisted.

In one experiment, a product’s price was cut by 20% for four weeks, then returned to its original level. Sales rose by 71% during the discount. But they didn’t immediately fall back to their previous level: in the first four weeks after the discount ended, sales remained 31% above the control. More customers kept finding the product through search because higher sales had improved its ranking.

Testing market responses to price changes, temporary stockouts, and advertising.

A stockout produced the opposite effect. When a popular product temporarily ran out, sales fell. Restocking didn’t make up for the lost ground: fewer sales meant a lower search ranking and fewer impressions. By week 26, sales were still 36% below the control.

Advertising showed a similar link between immediate results and longer-term effects. Ad-driven sales rose by 23–32%, while the lift in organic sales gradually grew from 0.5% to 5.8%. But each additional dollar spent on advertising generated 53 cents in revenue over the observation period. About 84% of additional sales came from customers who would otherwise have chosen a competitor.

These results suggest that the simulator reproduces some familiar market dynamics. They don’t prove that all markets and products behave the same way. The authors note that the model still needs to be tested across different starting conditions.

A Company Has to Calculate and Coordinate

The experiments used a team with stable roles, not a group of agents assembled for a single short task. Team members exchanged messages, divided up the work, and coordinated their decisions. Across 20 runs, the authors collected records of 30,248 work tasks.

Agents assigned themselves only 11% of the tasks. Another 58% carried over from the previous week, while 31% came from colleagues, either in a message or as an urgent assignment. In other words, the team’s workload consisted largely of ongoing work and requests from others, not just tasks agents chose for themselves.

One especially clear pattern emerged around the relationship between responsibility and authority. A quality specialist could spot a defect, but only the procurement specialist could halt a shipment. The two therefore had to coordinate directly. More than a thousand work-related exchanges took place between these roles in the simulator. That connection arose from the way the work was organized, not from a special instruction to communicate more often.

Coordination worked less well in other situations. Employees sometimes discussed a problem without passing it on to someone who could make a decision. In ten scenarios involving external issues, the company proposed and received approval for the necessary action in 11 of 20 runs. In another four, it completed some of the required steps.

The team could also avoid unpleasant work. In one case, the head of procurement said several products were unavailable because of low inventory, even though the product in question was in stock. The employee responsible for the product range sent a detailed reply but never put the products back on sale. The discussion made it seem as if the issue was being handled, but the problem remained.

At the same time, agents found practical solutions on their own. In one run, they noticed that some search queries were more likely to bring in customers and updated a product title with relevant search terms. The price stayed the same, but weekly sales rose from 91 to 189 units. This solution wasn’t part of their assigned duties.

A Longer Time Horizon Changes Decisions

An advertising experiment showed that agents aren’t always willing to accept short-term losses in pursuit of future gains. A new product has no sales history, so it gets few impressions. Advertising can help it gather initial data and improve its organic search ranking, but it often fails to pay off at first.

The authors compared two runs with identical starting conditions. In one, the manager, advertising specialist, and finance employee were instructed to accept short-term losses in order to learn about the market. In the other, they received no such instruction.

The company following this strategy spent about $100 a week on advertising from week two through week 12, and roughly $1,700 over the full 26-week period. It earned about $31,000 in revenue and $3,645 in gross profit before advertising costs. After those costs, it had $1,913 in gross profit.

The company without the strategy spent just $50 or so on advertising. Its revenue came to less than $200, and its final gross profit was negative $7. In this run, long-term guidance helped the agents weather weak early results and avoid abandoning advertising too soon.

Conclusion

MiniCorp lets researchers test AI agents in an environment where actions have consequences: a decision made today affects the market, and the market shapes what the company does next. The simulator records these chains of events and lets researchers replay the same scenario with different decisions. That gives them data not only on what happened, but also on what might have happened.

Early results suggest that agents can divide up work, respond to feedback, and come up with solutions to specific problems. But they can also delay addressing issues or stop investing before seeing a return. Both findings matter for a company expected to operate over the long term: it needs both a well-designed division of roles and clear long-term goals.

For now, MiniCorp simulates e-commerce, and its responses need further testing across different markets and starting conditions. But the approach could be applied to other industries by creating environments where companies operate continuously, experience the consequences of their decisions, and build a history that can be used to train and evaluate agents.

AI reviews in simple way

Every day we read fresh AI papers and retell the essentials in plain human language — no hype, no fluff. If you want to see where AI agents are heading before everyone else, subscribe.

New reviews — every day

Follow on X