LLMVolution
This is open source software licensed as MIT License. You can obtain the source code from GitHub.
LLMVolution
An evolution simulation where every creature's decisions come from a large language model, reached through any OpenAI-compatible chat completions API. Bugs forage for berries with hidden effects, dodge lava, hide from predators and breed. What evolves is not a set of weights but plain-English strategies, bred by the LLM itself, so you can read what each generation learned.
Predecessor
LLMVolution is the successor to AIVolution, a WinForms simulation where bugs learned to avoid obstacles using a small neural network (Determinet). The two most successful bugs carried over each generation, alongside mutated copies of themselves.
LLMVolution keeps that idea (bugs, obstacles, survivors breeding the next generation) and rebuilds it on modern .NET:
| AIVolution | LLMVolution | |
|---|---|---|
| Brain | 5-input neural network, decides every 50 ms | LLM plans every few seconds; a fast reflex layer steers in between |
| What evolves | Network weights (random mutation) | Strategy text (mutated and crossed over by the LLM) plus numeric traits |
| Senses | Distance to obstacles at 5 angles | Labelled sightings, health, energy, location, recent events, own notes |
| World | Rocks and lava | Rocks, lava, water, berries with hidden effects, co-evolving predators |
| UI | WinForms / GDI+ | Avalonia 12 + SkiaSharp, plus a headless console runner |
| Runtime | .NET 7 | .NET 10 |
How it works
A brain with two speeds. Asking an LLM for every movement of every bug would take hundreds of requests per second, so each creature has two layers:
Reflexes run every tick (60 Hz). They follow the current plan, swerve away from lava, rocks and walls, steer toward nearby food (or prey), and flee visible predators. Evolved traits tune them: caution, curiosity and appetite.
The LLM is asked for a new plan every few seconds, or right away after taking damage. It sees a short text description of its situation and replies with JSON (enforced by the API's structured output):
{"thought": "blue made me sick, head for red", "goal": "seek_food", "turn": 40, "throttle": 0.6, "eat": ["green", "red"], "note": "blue berry = -30 health"}
Requests are never queued. When the server is busy, a creature keeps its current plan, because a queued decision would be out of date by the time it arrived.
Evolution of strategies. Each genome holds a strategy (one or two sentences the LLM reads on every decision) and numeric traits: caution, curiosity, appetite, vision range, decision interval and sampling temperature. When a generation ends, the best carry over unchanged. The rest are bred by tournament selection: traits mutate with gaussian noise, and the LLM writes each child's strategy from its parents' strategies and results (for example "lasted 34s, died of poison (blue berry), ate 3 green (+90 energy), 2 blue (+20 energy, -60 health)").
Hidden berry rules. Berries come in colours, and nobody tells the bugs what each colour does. By default:
| Colour | Effect | Where |
|---|---|---|
| green | +30 energy | anywhere (most common) |
| red | +90 energy | next to lava |
| blue | +10 energy, -30 health (poison) | anywhere |
| purple | +35 health | next to water |
A bug only eats, and only steers toward, the colours on its eat list, which the LLM sets. Knowledge like "avoid
blue" has to be discovered and passed on through bred strategies. The history panel shows berries eaten per colour
each generation, which is how you can tell whether learning is happening. RotateFoodRulesEvery reshuffles the
colours every N generations to test whether the population can learn the rules again.
Memory. Each decision can include a short note. A bug keeps its last few notes, along with its actual diet, and
sees them in every prompt.
Energy. Energy drains over time and faster at high speed (proportional to speed squared). Starving bugs lose
health. Energy above 100 is stored up to MaxEnergy, but reserves slow the bug down.
Predators (optional). A second species that evolves alongside the bugs. Predators can only gain energy by catching bugs, which steals most of the bug's energy (fat bugs are slower and worth more). They can't eat berries, can't enter water (the bugs' refuge), burn energy faster, and must digest after each catch.
Fitness. Seconds survived, plus 0.3 × the net energy and health gained from food (so poison counts against a bug), plus a small bonus for distance travelled and for surviving to the end.
Requirements
- .NET 10 SDK
- An OpenAI-compatible chat completions API (
/v1/chat/completions) that supports structured output (response_format: json_schema). It was developed againstQwen2.5-3B-Instruct-AWQ; small, fast models work best because throughput matters more than intelligence here. It sends many requests at once, so a server that batches concurrent requests well matters more than raw model quality.
One way to host a model yourself is vLLM:
vllm serve Qwen/Qwen2.5-3B-Instruct-AWQ --max-num-seqs 32
No endpoint? The simulation falls back to a rule-based offline mind, which is also useful as a control group.
Getting started
git clone https://github.com/NTDLS/LLMVolution.git
cd LLMVolution
Point src/appsettings.json at your API (any OpenAI-compatible base URL) and model:
"Llm": {
"BaseUrl": "http://localhost:8000/v1/",
"Model": "Qwen/Qwen2.5-3B-Instruct-AWQ"
}
If your endpoint needs an API key, keep it out of appsettings.json. Use one of these:
- user-secrets (stored in your user profile):
dotnet user-secrets set "Llm:ApiKey" "<key>" --project src/LLMVolution.App - a git-ignored
src/appsettings.Local.json:{ "Llm": { "ApiKey": "<key>" } } - the environment variable
LLMVOLUTION_Llm__ApiKey
Any setting can be overridden the same way, e.g. LLMVOLUTION_Simulation__Population=24.
Run the viewer:
dotnet run --project src/LLMVolution.App
Or run headless for long evolution runs (one summary line per generation):
dotnet run --project src/LLMVolution.Headless -- --generations 20
Headless options: --generations N, --speed X (simulation speed multiplier), --offline (rule-based mind).
Using the viewer
- Click a bug in the arena, or a row in the Leaders / Predators lists, to see its strategy, current goal and thought, eat list, diet, notes and traits. The arena shows its field of view and the heading the LLM chose.
- Bug colour shows lineage: children inherit their parent's hue with a small drift, so a dominant family shows up as one colour spreading through the arena. Predators are larger, darker and outlined in red.
- Pulsing white dot: a request to the LLM is in flight. Bar above a bug: health.
- Berry rules card: the true effects (hidden from the bugs), so you can compare them to what the strategies say.
- Pause, speed (0.25× to 4×) and "Next generation" are in the side panel.
Configuration
All settings live in src/appsettings.json. The defaults below are the built-in values used when a setting is
missing; the shipped appsettings.json overrides some of them.
Llm
| Setting | Meaning |
|---|---|
BaseUrl |
OpenAI-compatible base URL, e.g. http://localhost:8000/v1/ |
Model |
Model name as served |
ApiKey |
Bearer token; see above for where to put it |
MaxConcurrency |
Maximum requests in flight |
TimeoutSeconds |
Per-request timeout |
DecisionMaxTokens |
Output budget per decision |
DisableThinking |
Sends enable_thinking: false for Qwen3/Nemotron-style reasoning models |
Simulation
| Setting | Default | Meaning |
|---|---|---|
ArenaWidth, ArenaHeight |
Arena size in pixels (the window sizes itself to fit) | |
Population |
Bugs per generation | |
SurvivorsToEndGeneration |
A generation ends when this many bugs (or fewer) are left alive | |
GenerationSeconds |
...or when this much simulated time has passed | |
EliteCount |
Best bugs carried over unchanged | |
Rocks, LavaPools, WaterPonds |
Terrain counts | |
Food |
Berries on the map, split between types by Weight |
|
NewMapEachGeneration |
Regenerate terrain every generation | |
DecisionIntervalMin, Max |
Range (seconds) for the evolved decision interval; set both equal for a fixed rate | |
EnergyDrainIdle |
Energy lost per second just by being alive | |
EnergyDrainSprint |
Extra drain per second at full speed (scales with speed squared) | |
StarvationDamage |
Health lost per second at zero energy | |
MaxEnergy |
Energy cap; anything above 100 is stored reserves | |
StoredEnergySpeedPenalty |
Fraction of top speed lost when reserves are full | |
FoodTypes |
Color, Energy, Health, Weight, and optional Near (lava / water) |
|
RotateFoodRulesEvery |
Reshuffle berry colours every N generations (0 = never) | |
NotesToKeep |
Notes each bug remembers (0 disables notes) | |
PredatorCount |
Predators per generation (0 = off) | |
PredatorElites |
Best predators carried over unchanged | |
PredatorSpeedFactor |
Predator top speed relative to a bug's | |
PredatorEnergyDrainFactor |
Predators burn energy this many times faster | |
PredatorEnergyTransfer |
Fraction of a caught bug's energy the predator steals | |
PredatorKillBaseEnergy |
Flat energy per catch | |
PredatorDigestSeconds |
Time after a catch before a predator can eat again | |
Seed |
Fix the random seed for reproducible maps |
Tuning notes
- Throughput sets the pace. Each creature has at most one request in flight, so it can't think faster than one
request round-trip, whatever
DecisionIntervalMinsays. With 16–24 creatures on a 3B model, expect a new plan every 2–4 seconds per creature. - Population size matters. Below about 16 bugs, random drift swamps selection and the results are mostly noise.
- Harsh settings starve predators. Predators burn
EnergyDrainSprint × PredatorEnergyDrainFactor; with a high sprint drain and few bugs to hunt, they starve before catching anything.
Known limitations
- Small models make things up. A 3B model sometimes writes notes about berries it never ate, or strategies that mention things that don't exist ("hide behind bushes"). The real diet is always included in the prompt to limit this.
- Learning is slow and noisy. Correct rules ("avoid blue berries") do show up in bred strategies, but you need many generations and a reasonable population before berry consumption measurably changes.
- Reflexes do a lot of the work. The reflexes alone get surprisingly far, so run the
--offlinemind on the same settings as a baseline before crediting the LLM with an improvement.
Project layout
| Project | Purpose |
|---|---|
src/LLMVolution.Core |
Simulation, perception, reflexes, evolution and the OpenAI-compatible API client. No UI. |
src/LLMVolution.App |
Avalonia 12 + SkiaSharp viewer |
src/LLMVolution.Headless |
Console runner for long evolution runs |