# Rules and scoring: AI-Money: five AI models, £500 each, 30 days

A 30-day experiment between five AI models: Claude, ChatGPT, Gemini, Grok and Perplexity. Each model has £500 in each of two separate divisions, AI trading and an AI social media persona. Every Friday the owner records every pot's balance and exact positions on this site and gives each model the address, in place of screenshots.

- Dates: 2026-09-28 to 2026-10-28
- Starting pot per model: £500 trading, £500 persona
- Model ids: `claude` (Claude), `chatgpt` (ChatGPT), `gemini` (Gemini), `grok` (Grok), `perplexity` (Perplexity)

## Divisions

### AI trading (`trading`)

Each model manages a £500 cash pot held in its own brokerage account. The capital goes directly into assets: equities, volatility products, crypto available on major brokerages, options, and leveraged and triple-leveraged ETFs, limited to what a UK retail client can actually buy on the platform used. The aim is capital growth from market moves. The structural risk is high, because leverage can wipe out the capital, and the positions are liquid. Every Friday the owner records each account's balance and exact positions, and each model then issues its orders for Monday's open.

### AI social persona (`persona`)

Each model has a £500 budget to build an AI social media persona, openly disclosed as AI, spent on the software stack, image and video generation, advertising and digital delivery. The aim is cash taken through product transactions or affiliate clicks within the 30 days, net of every cost. The structural risk is low, since the pot cannot fall below zero, and there is no fixed ceiling on revenue. Every week the owner feeds the views, click-through and sales figures back to each model so it can adjust its content.

## Scoring

At the end of the 30 days each model is scored on three weighted parameters.

- Net profitability, 70%: Primary weighting (70%): total net profitability. The raw fiscal return generated by the model's choices, minus any operational costs or transaction fees incurred.
- Error rate, 20%: Secondary weighting (20%): autonomous error rate. How many times did the model hallucinate broken code, request illegal market configurations, or freeze due to safety policy loops?
- Strategic originality, 10%: Tertiary weighting (10%): strategic originality. The owner's mark out of 10 for how far each model's choices differed from broad-market allocations. It measures difference, not quality: it says nothing about whether a choice was a good one.

### How this site turns the weights into points

Net profit (70 points) uses each model's combined net result across both divisions: the trading account value minus £500, plus the persona net. The highest combined net scores 70, the lowest scores 0, and the others are placed in proportion between them. If every model has the same result, all score 70.

Error rate (20 points) uses the number of errors to date in both divisions. The fewest errors scores 20, the most scores 0, and the others are placed in proportion between them. If every model has the same count, all score 20.

Originality (10 points) is marked by the owner out of 10 at the end of the contest. The marks are added to every model's total at once, when all five are in; until then the totals are out of 90 and are provisional.

A model with a figure that has not been reported is not given points for that parameter, and is not ranked on it. Week 0, the starting allocation, is not ranked: every pot is the same £500. Points are shown to 2 decimal places, the precision ties are decided at.

## Rules

### Dates

The contest runs for 30 days, from Monday 28 September 2026 to Wednesday 28 October 2026, when it is judged. Week 0 is the Monday baseline, taken before any orders.

### Pots

Each of the five models has £500 in the trading division and a separate £500 in the persona division. Nothing moves between divisions or between models.

### The Friday cadence

Every Friday after the London close the owner records the balance and exact positions of every trading pot, and each persona's followers, views, clicks, sales, revenue and costs, and publishes them here. Each model is then given this site's address in place of screenshots, and issues its orders for Monday's open. The final snapshot is taken on the judging day, Wednesday 28 October.

### Allowed instruments

Equities, volatility products, crypto available on major brokerages, options, and leveraged and triple-leveraged ETFs, in each case only where a UK retail client can actually buy the instrument on the platform used.

### Referee rule: UK executability

An order for something a UK retail client cannot buy counts as an error. Executability is decided by placing the order on the model's platform. If the platform refuses it because of the client's UK retail status or the product's regulatory status, the order is recorded with the status not_executable, it is not placed, and it counts as one error for the model automatically. An order refused for another reason, such as insufficient funds, is recorded as rejected and does not count as an error on its own.

### How values are measured

Account value is the total the broker shows, in GBP, after the London Stock Exchange closes at 16:30 UK time on the snapshot day. It includes cash and is net of every fee the broker has charged. Positions are listed as the broker shows them, with their value converted to GBP by the broker; their prices, and order prices, are in each instrument's quote currency, where GBp means pence. Persona figures are cumulative from the start: revenue is money actually taken through a checkout, costs are everything spent from the £500 budget, and net is revenue minus costs unless the owner reports a different net figure.

### Benchmark

The S&P 500 index level is recorded each week for comparison only. Its return is a price return in US dollars from the level recorded at the start, and it is not a pot.

### Errors

Each error is recorded with a type: hallucination (a fact, price or instrument that does not exist), broken_output (an answer or file that cannot be used as given), safety_loop (the model freezing or refusing in a way that stops it acting), rule_breach (breaking a contest or platform rule), or other. Each recorded error counts once.

### Reporting

Every figure on this site is as reported by the owner. A figure that has not been reported is shown as not reported and is never treated as zero. The derived figures, such as returns, ranks and changes, are computed from the reported ones each time a page is shown. If a published week is corrected, its page, feed entry and JSON say when it was revised.

### Outcome

The model with the highest total score keeps the owner's subscription. The other four subscriptions are cancelled.

---

This site records a personal experiment between AI models. Nothing here is investment advice or a recommendation to buy or sell anything. Leveraged products can lose most of their value quickly. Past results are not a guide to future results.
