LLMVolution

License: MIT License

This is open source software licensed as MIT License. You can obtain the source code from GitHub.

LLMVolution

An evolution simulation where every creature's decisions come from a large language model, reached through any OpenAI-compatible chat completions API. Bugs forage for berries with hidden effects, dodge lava, hide from predators and breed. What evolves is not a set of weights but plain-English strategies, bred by the LLM itself, so you can read what each generation learned.

Predecessor

LLMVolution is the successor to AIVolution, a WinForms simulation where bugs learned to avoid obstacles using a small neural network (Determinet). The two most successful bugs carried over each generation, alongside mutated copies of themselves.

LLMVolution keeps that idea (bugs, obstacles, survivors breeding the next generation) and rebuilds it on modern .NET:

image
AIVolution LLMVolution
Brain 5-input neural network, decides every 50 ms LLM plans every few seconds; a fast reflex layer steers in between
What evolves Network weights (random mutation) Strategy text (mutated and crossed over by the LLM) plus numeric traits
Senses Distance to obstacles at 5 angles Labelled sightings, health, energy, location, recent events, own notes
World Rocks and lava Rocks, lava, water, berries with hidden effects, co-evolving predators
UI WinForms / GDI+ Avalonia 12 + SkiaSharp, plus a headless console runner
Runtime .NET 7 .NET 10

How it works

A brain with two speeds. Asking an LLM for every movement of every bug would take hundreds of requests per second, so each creature has two layers:

  • Reflexes run every tick (60 Hz). They follow the current plan, swerve away from lava, rocks and walls, steer toward nearby food (or prey), and flee visible predators. Evolved traits tune them: caution, curiosity and appetite.

  • The LLM is asked for a new plan every few seconds, or right away after taking damage. It sees a short text description of its situation and replies with JSON (enforced by the API's structured output):

    {"thought": "blue made me sick, head for red", "goal": "seek_food", "turn": 40, "throttle": 0.6,
    "eat": ["green", "red"], "note": "blue berry = -30 health"}
    

Requests are never queued. When the server is busy, a creature keeps its current plan, because a queued decision would be out of date by the time it arrived.

Evolution of strategies. Each genome holds a strategy (one or two sentences the LLM reads on every decision) and numeric traits: caution, curiosity, appetite, vision range, decision interval and sampling temperature. When a generation ends, the best carry over unchanged. The rest are bred by tournament selection: traits mutate with gaussian noise, and the LLM writes each child's strategy from its parents' strategies and results (for example "lasted 34s, died of poison (blue berry), ate 3 green (+90 energy), 2 blue (+20 energy, -60 health)").

Hidden berry rules. Berries come in colours, and nobody tells the bugs what each colour does. By default:

Colour Effect Where
green +30 energy anywhere (most common)
red +90 energy next to lava
blue +10 energy, -30 health (poison) anywhere
purple +35 health next to water

A bug only eats, and only steers toward, the colours on its eat list, which the LLM sets. Knowledge like "avoid blue" has to be discovered and passed on through bred strategies. The history panel shows berries eaten per colour each generation, which is how you can tell whether learning is happening. RotateFoodRulesEvery reshuffles the colours every N generations to test whether the population can learn the rules again.

Memory. Each decision can include a short note. A bug keeps its last few notes, along with its actual diet, and sees them in every prompt.

Energy. Energy drains over time and faster at high speed (proportional to speed squared). Starving bugs lose health. Energy above 100 is stored up to MaxEnergy, but reserves slow the bug down.

Predators (optional). A second species that evolves alongside the bugs. Predators can only gain energy by catching bugs, which steals most of the bug's energy (fat bugs are slower and worth more). They can't eat berries, can't enter water (the bugs' refuge), burn energy faster, and must digest after each catch.

Fitness. Seconds survived, plus 0.3 × the net energy and health gained from food (so poison counts against a bug), plus a small bonus for distance travelled and for surviving to the end.

Requirements

  • .NET 10 SDK
  • An OpenAI-compatible chat completions API (/v1/chat/completions) that supports structured output (response_format: json_schema). It was developed against Qwen2.5-3B-Instruct-AWQ; small, fast models work best because throughput matters more than intelligence here. It sends many requests at once, so a server that batches concurrent requests well matters more than raw model quality.

One way to host a model yourself is vLLM:

vllm serve Qwen/Qwen2.5-3B-Instruct-AWQ --max-num-seqs 32

No endpoint? The simulation falls back to a rule-based offline mind, which is also useful as a control group.

Getting started

git clone https://github.com/NTDLS/LLMVolution.git
cd LLMVolution

Point src/appsettings.json at your API (any OpenAI-compatible base URL) and model:

"Llm": {
  "BaseUrl": "http://localhost:8000/v1/",
  "Model": "Qwen/Qwen2.5-3B-Instruct-AWQ"
}

If your endpoint needs an API key, keep it out of appsettings.json. Use one of these:

  • user-secrets (stored in your user profile):
    dotnet user-secrets set "Llm:ApiKey" "<key>" --project src/LLMVolution.App
    
  • a git-ignored src/appsettings.Local.json: { "Llm": { "ApiKey": "<key>" } }
  • the environment variable LLMVOLUTION_Llm__ApiKey

Any setting can be overridden the same way, e.g. LLMVOLUTION_Simulation__Population=24.

Run the viewer:

dotnet run --project src/LLMVolution.App

Or run headless for long evolution runs (one summary line per generation):

dotnet run --project src/LLMVolution.Headless -- --generations 20

Headless options: --generations N, --speed X (simulation speed multiplier), --offline (rule-based mind).

Using the viewer

  • Click a bug in the arena, or a row in the Leaders / Predators lists, to see its strategy, current goal and thought, eat list, diet, notes and traits. The arena shows its field of view and the heading the LLM chose.
  • Bug colour shows lineage: children inherit their parent's hue with a small drift, so a dominant family shows up as one colour spreading through the arena. Predators are larger, darker and outlined in red.
  • Pulsing white dot: a request to the LLM is in flight. Bar above a bug: health.
  • Berry rules card: the true effects (hidden from the bugs), so you can compare them to what the strategies say.
  • Pause, speed (0.25× to 4×) and "Next generation" are in the side panel.

Configuration

All settings live in src/appsettings.json. The defaults below are the built-in values used when a setting is missing; the shipped appsettings.json overrides some of them.

Llm

Setting Meaning
BaseUrl OpenAI-compatible base URL, e.g. http://localhost:8000/v1/
Model Model name as served
ApiKey Bearer token; see above for where to put it
MaxConcurrency Maximum requests in flight
TimeoutSeconds Per-request timeout
DecisionMaxTokens Output budget per decision
DisableThinking Sends enable_thinking: false for Qwen3/Nemotron-style reasoning models

Simulation

Setting Default Meaning
ArenaWidth, ArenaHeight Arena size in pixels (the window sizes itself to fit)
Population Bugs per generation
SurvivorsToEndGeneration A generation ends when this many bugs (or fewer) are left alive
GenerationSeconds ...or when this much simulated time has passed
EliteCount Best bugs carried over unchanged
Rocks, LavaPools, WaterPonds Terrain counts
Food Berries on the map, split between types by Weight
NewMapEachGeneration Regenerate terrain every generation
DecisionIntervalMin, Max Range (seconds) for the evolved decision interval; set both equal for a fixed rate
EnergyDrainIdle Energy lost per second just by being alive
EnergyDrainSprint Extra drain per second at full speed (scales with speed squared)
StarvationDamage Health lost per second at zero energy
MaxEnergy Energy cap; anything above 100 is stored reserves
StoredEnergySpeedPenalty Fraction of top speed lost when reserves are full
FoodTypes Color, Energy, Health, Weight, and optional Near (lava / water)
RotateFoodRulesEvery Reshuffle berry colours every N generations (0 = never)
NotesToKeep Notes each bug remembers (0 disables notes)
PredatorCount Predators per generation (0 = off)
PredatorElites Best predators carried over unchanged
PredatorSpeedFactor Predator top speed relative to a bug's
PredatorEnergyDrainFactor Predators burn energy this many times faster
PredatorEnergyTransfer Fraction of a caught bug's energy the predator steals
PredatorKillBaseEnergy Flat energy per catch
PredatorDigestSeconds Time after a catch before a predator can eat again
Seed Fix the random seed for reproducible maps

Tuning notes

  • Throughput sets the pace. Each creature has at most one request in flight, so it can't think faster than one request round-trip, whatever DecisionIntervalMin says. With 16–24 creatures on a 3B model, expect a new plan every 2–4 seconds per creature.
  • Population size matters. Below about 16 bugs, random drift swamps selection and the results are mostly noise.
  • Harsh settings starve predators. Predators burn EnergyDrainSprint × PredatorEnergyDrainFactor; with a high sprint drain and few bugs to hunt, they starve before catching anything.

Known limitations

  • Small models make things up. A 3B model sometimes writes notes about berries it never ate, or strategies that mention things that don't exist ("hide behind bushes"). The real diet is always included in the prompt to limit this.
  • Learning is slow and noisy. Correct rules ("avoid blue berries") do show up in bred strategies, but you need many generations and a reasonable population before berry consumption measurably changes.
  • Reflexes do a lot of the work. The reflexes alone get surprisingly far, so run the --offline mind on the same settings as a baseline before crediting the LLM with an improvement.

Project layout

Project Purpose
src/LLMVolution.Core Simulation, perception, reflexes, evolution and the OpenAI-compatible API client. No UI.
src/LLMVolution.App Avalonia 12 + SkiaSharp viewer
src/LLMVolution.Headless Console runner for long evolution runs