This series builds Westworld of Warcraft: a private World of Warcraft server populated by machine-controlled characters that quest, group, raid, trade, and talk. It covers the injection and protocol runtimes, the behavior hierarchy, navigation and physics, the personality and storyline layers, the social fabric and economy, the advisory AI boundary, and the evidence that proves the world is actually alive.
Jared Rhodes
· September 14, 2026–September 28, 2026
· 3 posts
https://jaredrhodes.com
Westworld of Warcraft: A Server That Plays Itself
September 14, 2026
In 2005 I played a hunter who would not group with anyone who typed in all caps. I never learned his name. I remember the rule.
That is the thing worth building here. Not a bot that clears a dungeon - plenty of those exist. A population. A server where the other characters have habits, prices, grudges, and schedules, and where a human logging in at two in the morning can find five of them willing to run Wailing Caverns.
Westworld of Warcraft is that server. This series is how it is built and why each part is shaped the way it is.
Why One Bot Was Never the Point
“Make bots that play WoW” is a solved and boring problem. “Make a server that behaves like it has a thousand people on it” is neither.
The gap between those two sentences is where all the engineering lives:
One bot grinding boars is a state machine. A thousand bots grinding boars is a scheduling and economy problem.
One bot pathing to a vendor is A*. A thousand bots pathing across Azeroth is a mesh generation and caching problem.
One bot casting Fireball is a rotation. A thousand bots casting Fireball with identical timing is a detectable signal that no human population produces.
One bot answering a whisper is a template. A thousand bots answering whispers is a rate-limiting and anti-griefing problem.
I keep one north star for the whole project: a new human player logging in cannot tell, from gameplay observation alone, that the population is machine-controlled.
What “Alive” Actually Means
Vague goals produce vague systems, so before writing anything I made myself define “alive” as four testable properties:
Always-available activities. Any legal activity - questing zone, dungeon, raid, battleground, profession route, world event, world boss - has participants available inside the activity’s normal group-form window. A human request for a level-appropriate dungeon group resolves in under five minutes.
Continuous progression. Bots not serving a human request advance toward their own roster goals: level, gear, attunement, reputation, profession, gold, mount, PvP rank. An idle bot is a bug.
A living economy. The auction house shows posting and bidding activity around the clock. Vendors see traffic. Banks see deposits. Mail moves. The economy reaches steady state without seeding.
Operator clarity. The console shows me the highest-volume errors, the active activities, the population, and the scaling pressure points. Choosing my next engineering task is a query, not a guess.
Notice what is not on that list. There is no requirement that bots fake incompetence - misspelled trade chat, aimless wandering, theater. Social texture itself is absolutely in scope (the storyline and social-fabric posts depend on it); what is out of scope is performing incompetence to seem human. Indistinguishability here is a property of timing and routing.
Three Thousand Characters, On Purpose
The target population is 3,000 characters, and I picked its distribution deliberately.
Those 3,000 are not all logged in at once. Every character carries its own online window, so a typical hour has roughly a thousand of them in the world. Where this series does per-hour arithmetic - trade posts, mail volume, chat budgets - it runs at a thousand bots and means that number. The 3,000 is the roster; the thousand is the crowd.
Dimension
Target
Faction split
Roughly 50/50, configurable per realm
Class and spec coverage
Every race, class, and spec combination present at every five-level bracket, weighted toward level 60
Professions
Every primary profession represented at 300 skill; cooking, first aid, and fishing universal
PvP
Enough queue depth for Warsong Gulch at 10v10, Arathi Basin at 15v15, and Alterac Valley at 40v40 inside bracket boundaries
Raids
Enough attuned level 60s for one concurrent raid in each tier: Onyxia, Molten Core, Blackwing Lair, Zul’Gurub, AQ20, AQ40, Naxxramas
A RosterPlanner owns account-level decisions and enforces the coverage rules, in order. The counts below are the vanilla 1.12.1 roster; the later clients add a class and professions to both.
Faction bootstrap. If the plan needs a shaman and the account has no Horde characters, a Horde character gets created first.
Class coverage. All nine classes reach 60 before any class is duplicated at 60.
Profession coverage. All nine primary professions distributed; none left unrepresented at 300.
Spec diversity. Each class fields at least one of each role it can fill.
PvP rank. The roster holds characters at each rank band Alterac Valley objectives need.
This is why the server does not end up as 3,000 fury warriors. Somebody has to be the enchanter.
What Is Actually Running
The stack reads from the bottom up. Exports/GameData.Core holds the game interfaces and shared contracts with zero dependencies - it is the bottom of the stack on purpose, because everything else stands on it. Exports/BotCommLayer carries protobuf over TCP with length framing, along with the .proto sources and their generated C#. Exports/BotRunner is the behavior engine itself: task stack, objective decomposition, shared by both runtimes. Exports/WoWSharpClient is a pure C# implementation of the WoW protocol - packets, opcodes, auth, movement. Exports/Navigation wraps C++ Detour pathfinding plus the PhysicsEngine.dll collision target, and Exports/Loader with Exports/FastCall provide the C++ CLR-injection bootstrap and structured-exception-wrapped fast calls.
The service tier does the coordinating. Services/WoWStateManager is the orchestrator: bot lifecycle, foreground injection, IPC listeners, activity registry, legality. Services/PathfindingService runs A* routes over the native navigation layer, and Services/SceneDataService feeds collision and scene geometry to background bots. Services/DecisionEngineService produces advisory recommendations for objectives, rewards, rotations, threat, chat, and personality; Services/PromptHandlingService hosts the persona and dialogue runtime plus the storyline graph store. BotProfiles carries the per class and spec combat rotations.
On top sit the loopback-only Blazor Server consoles, UI/OperatorConsole and UI/StorylineManager, on ports 5167 and 5157, with the storyline runtime API alongside them on 5147, and UI/Systems/Systems.AppHost, which does the .NET Aspire orchestration of the Docker stack and the services.
The dependency direction is strict, and a test enforces it:
Interfaces live in the lower layers. Implementations live in the higher ones. An Exports/ project may never reference Services/, UI/, or Tests/. When someone tries, ProjectLayeringTests fails the build.
Two Ways to Drive a Character
There are exactly two runtimes, they share the behavior engine and the game interfaces, and neither is a fallback for the other.
The foreground runtime injects a native loader into a real WoW.exe, bootstraps the .NET runtime in-process, and drives the character through direct memory reads and writes plus Lua. You reach for it when you need true client parity: rendering, exact physics, packet captures that serve as parity baselines.
The background runtime is headless. A pure C# implementation of the WoW protocol connects to the world server with no game client at all, which is how you get many bots cheaply, or CI without a GUI.
Both are real, both ship. The next post takes them apart.
Why a Legacy Private Server
This question comes up immediately and it deserves a direct answer. The supported clients are Vanilla 1.12.1, Burning Crusade 2.4.3, and Wrath of the Lich King 3.3.5a, running against a locally hosted MaNGOS-family world server. Modern retail WoW is not supported and is not a goal.
Four reasons, in roughly this order.
Consent. Everyone on the server is either the operator or something the operator started. Nobody’s competitive experience is being degraded.
Stability. 1.12.1 is a fixed target. Memory offsets, opcodes, and physics constants do not move under you between patches.
Observability. The operator owns the world database, the server logs, and the SOAP command interface. You can ask the world questions directly.
Research value. The interesting problems - coordination, economy, planning, indistinguishability - do not require the newest client to be interesting.
The stated purpose is intellectual exploration. That is a stronger position when the environment is one you own outright.
The Invariants
These survive every refactor. Breaking one is a priority-zero bug, and most of the rest of this series is downstream of them.
Invariant
Why
No blind sequences
Counters, sleeps, and fixed repeat-N-times loops are banned for state validation. Gate on memory, packet, snapshot, or explicit API state. A bot that waits three seconds and hopes is not a bot, it is a superstition.
Foreground is ground truth
When the foreground and background runtimes disagree about physics or movement, the foreground is right and the background is wrong.
StateManager owns orchestration
Tests, the UI, and external callers never bypass it to talk to a bot directly.
Geometry has exactly two owners
The pathfinding and scene-data services answer every world-geometry question. Bot code does not load map tiles.
Tests assert through snapshots
A test that reaches into internal bot state instead of reading a published snapshot is testing the wrong thing.
No skipping for “resource not found”
If a fishing pool exists in the world database and the bot cannot find it, that is a detection or pathfinding bug, not a reason to skip the test. Walk further.
The catalog drives legality
Every rejection of an illegal activity cites a specific catalog field. No ad-hoc legality logic inside the behavior engine.
The one that changes the most code is the first. “No blind sequences” is why every task in the system carries a verification predicate, and why the behavior hierarchy in Activity, Objective, Task, Action looks the way it does.
Where This Series Goes
Four groups. Foundations covers how a character gets driven, how behavior is structured, and how movement is made correct. Population covers where personality comes from, who the standing cast is, how they talk, and how the economy forms. World and intelligence covers server-wide time, the boundary around machine learning and language models, and how a human asks the world for something. Proof covers how you test a world, and the synthesized reference architecture.
The population posts are the ones I would read first if I were you. Architecture is the substrate; the cast is the point.
The Idle Bot
Every failure worth designing against on this project is loud except one, and the quiet one is the idle bot.
A level 34 rogue completes its zone quest chain in Stranglethorn Vale.
The progression planner has no next objective because the bracket’s catalog rows are incomplete.
The bot stands in Booty Bay.
Nothing errors. Nothing alerts. The dashboard is green.
That is worse than a crash. A crash tells you where it hurts; a standing rogue tells you nothing until a human wanders past and sees a statue. That is why “an idle bot is a bug” sits in the definition of alive at the top of this post, and why the operator console surfaces population activity distribution as a first-class panel.
How I Will Know It Worked
Most of the finish line is unglamorous, and I expect to hit it quietly. Every activity in the catalog gets an automated test that drives a request through to a real group and a real completion. All 27 class and spec combat profiles pass live validation in both runtimes. The operator console renders population, active activities, top errors, and queue depth, and picks up config changes without a restart. Reproducible client crashes get a hardening fix or a written mitigation. None of that is interesting once it passes; it is only interesting while it fails.
Three of the criteria carry an actual argument. The staged load run has to reach 3,000 concurrent bots with snapshot latency holding under half a second at the ninety-ninth percentile, because the population is the product and a design that only works at fifty bots is a different design. Normal-operation logs have to be quiet enough that any warning is signal, because the idle bot above is invisible inside a noisy log. And every pattern that landed has to be written down as a technique someone could reuse, with an eye toward whether it transfers past this one game. That last one matters more than it looks. If none of this transfers to another game, then what got built is a WoW bot, not a method.
There are two ways to make a character in a virtual world do something. You can move the hands that hold the controller, or you can be the controller.
Westworld of Warcraft does both, on purpose, and the tension between them is the most productive constraint in the codebase.
Pick One and You Lose Something
Inject into the real client only. Every bot needs a WoW.exe process, a window, a GPU context, and roughly a gigabyte of address space. Thirty bots is a heroic machine. Three thousand is a data center. Continuous integration on a headless Linux runner is off the table permanently.
Emulate the protocol only. Now a bot is cheap - hundreds per machine, no GUI, trivially scriptable in CI. But you have inherited the entire client. Movement physics, collision, transport state, spell timing, update-field semantics: all of it now lives in code you wrote, and the only way to know whether you got it right is to compare against the thing you were trying to avoid running.
So both ship. The foreground runtime is the oracle. The background runtime is the fleet.
One Engine, Two Bodies
Foreground
Background
Process
Inside a live WoW.exe
Its own headless process
Game state
Direct memory reads and writes, plus the client’s own Lua
Parsed from SMSG_* packets into an object manager
Movement
The client’s real physics
PhysicsEngine.dll, ported from the client binary
Cost per bot
One full game client
A socket and a state machine
Runs in CI
No
Yes
Authority
Ground truth
Must match ground truth
What they share is everything above the seam: GameData.Core interfaces, the BotRunner behavior engine, the class and spec rotation profiles, and the protobuf transport. A task like GoToTask has no idea which runtime it is executing in. That is the whole design goal - one behavior engine, two ways of reaching the world.
Getting Inside the Client
The foreground path is process injection with a .NET twist, and the twist is the interesting part.
The host side is conventional Windows work:
WoWStateManager launches or locates the client process.
It waits for a real window and a real world state - no fixed sleeps, per the no-blind-sequences rule.
OpenProcess, then allocate memory inside the target.
Write the path to Loader.dll into that allocation.
Point CreateRemoteThread at LoadLibrary with that path.
The in-process side is where .NET 8 changes the old recipe. The classic injection tutorial uses mscoree.dll and the .NET Framework hosting API. That API still exists on Windows; what it cannot do is host .NET 8. Loader.dll instead uses hostfxr:
Loader.dll entry point
-> resolve hostfxr via nethost
-> hostfxr_initialize_for_runtime_config(ForegroundBotRunner.runtimeconfig.json)
-> get_function_pointer(load_assembly_and_get_function_pointer)
-> load ForegroundBotRunner.dll
-> invoke ForegroundBotRunner.Loader::Load
Three consequences fall out of that, and each one bit me before it got written down:
A runtimeconfig.json is mandatory. Hosting a modern runtime means initializing it from a declared configuration. Ship the config next to the assembly or nothing happens.
The entry point signature is fixed. A static method with the exact expected shape. Get it wrong and you get a null function pointer with no diagnostic.
Bitness is not negotiable. The 1.12.1 client is 32-bit, so Loader and FastCall build as x86. The native navigation and physics library builds x64 because it lives in the services. Two toolchains, one solution, and the build script tells you which one is missing.
Bootstrap runs on its own thread rather than in DllMain, because doing real work under the loader lock is how you deadlock a game client. A shutdown event is signaled on process detach so teardown is deterministic.
Once managed code is live inside the process, the bot has what no protocol client can have: the client’s own object manager, the client’s own Lua state, and the client’s own physics already computed. It reads the player’s position out of memory rather than deriving it, and calls game functions directly through structured-exception-wrapped thunks so a bad call surfaces as an error instead of taking the process with it.
Warden, the legacy anti-cheat, is disabled on injection. On a private research server with no competitive stake this is housekeeping, not evasion - the alternative is the client terminating itself mid-experiment.
Two Operational Rules Learned the Hard Way
Kill the client before you build. The injector loads native DLLs from the build output directory. A running client holds a lock on them, and MSBuild reports it as a file-copy error that looks nothing like the actual cause. I lost more time to that one message than I care to admit.
Version the offsets. Memory offsets are specific to exact client builds - 1.12.1 build 5875, 2.4.3 build 8606, 3.3.5a build 12340. A bot that logs in and then does nothing intelligible is almost always a client-build mismatch.
Being the Client Instead
The background runtime never touches a game client. WoWSharpClient is a pure C# implementation of the wire protocol: well over a hundred distinct opcodes handled in each direction.
Here is the stack it has to reproduce:
Auth - SRP6 challenge and proof against the realm daemon, then the realm list.
World handshake - session-key proof and header encryption on the world connection.
Object updates - parse SMSG_UPDATE_OBJECT and its update masks into a live object graph.
Movement - emit MSG_MOVE_* with correct flags, and pair server acknowledgements with the state transitions that caused them.
Transport - track boat, zeppelin, and elevator state so a bot on a moving object has coherent coordinates.
The movement layer is where the difficulty concentrates, because movement is the one thing the server actively checks. The parity contract is explicit:
The background runtime sends the same opcode, at the same flag state, with the same payload the real client would send.
Timing tolerance is +/-100 ms for self-initiated movement and +/-10 ms for server-initiated movement such as a forced root, a forced speed change, or a teleport.
Server acknowledgements are paired by opcode and state transition. A mismatched acknowledgement is a bug.
Ten milliseconds for server-initiated movement sounds severe until you watch a bot get rooted and answer with the wrong acknowledgement. The server’s correction fights the client’s state, and the character stutters in place like a bad connection. Which, from the server’s point of view, is exactly what it is.
The Parity Discipline
Two implementations of the same behavior will drift. The only question is whether you find out from a test or from a screenshot.
The rule is short: when the runtimes disagree, the foreground is right. The background implementation is a port of the client’s behavior, so a difference is by definition a porting defect.
That rule needs teeth, and the teeth are what counts as proof. A parity row closes on decompilation evidence naming a specific routine in the client binary, plus a canary that fails before the fix and passes after it. A recorded session that replays without visible error, a test asserting that a named route completes, and a screenshot of a bot standing in the right place can all be true while the port is still wrong. The split gets its row-by-row treatment in Teaching a Bot to Walk (coming soon), because the physics port is where it has to hold.
Packet capture and movement recording from the foreground runtime are still valuable, as observation. A capture tells you something differs. It does not tell you which routine differs, and it cannot close an implementation row on its own.
This distinction is the difference between a port that converges and a port that oscillates forever. Replay harnesses feel like proof because they are red and then green. They are actually a very expensive way to notice that something changed.
The Slope That Looked Like Success
The divergence that made me write the parity rule down never raised anything: a background bot climbed a Redridge slope the real client refuses, arrived, verified its task, and left a clean snapshot, and the disagreement only surfaced weeks later when a foreground bot on the same route took the long way around and blew a group-form timeout. I tell that story properly in Teaching a Bot to Walk (coming soon), where the physics argument lives.
What it settles here is the direction of blame. The cheap runtime was wrong in a way that looked like success, so widening the timeout would have buried the only signal I had. The disagreement itself is the artifact worth keeping: reduce it to a slope-threshold canary against the client’s own collision behavior, fix the native side, and let the timeout stand.
Most of the operational rules on this project fall out of that same instinct. When the background moves where the foreground cannot, the fix goes into the background physics rather than into the mesh that would hide it. When foreground offsets stop working, confirm the exact client build before touching a line of logic. When security software blocks injection, allow the build output rather than weakening the loader. When the native DLL copy fails during a build, find the specific client process holding the lock and kill that one. And when the background’s acknowledgement stops matching after a forced root, it is a movement-protocol bug, not server flakiness.
When the Two Bodies Agree
Most of what I check is plumbing. A behavior task compiles and runs unchanged in both bodies. Foreground injection is reliable against all three supported client builds. A background bot logs in, enters the world, moves, fights, and logs out with no client present, and its movement holds the packet parity contract inside the stated timing tolerances. The whole background suite runs in continuous integration with no GUI anywhere near it.
Two of the criteria are the ones that decide whether the oracle is worth having. Every closed physics parity row has to cite a specific routine in the client binary and carry a canary that fails without the fix, because a row closed on a passing replay is a row that will quietly reopen. And a disagreement between the runtimes always opens a defect against the background implementation. That is a promise about where blame goes, and keeping it is what makes the foreground worth trusting.
Related Posts
The behavior engine both runtimes share is the subject of Activity, Objective, Task, Action. The physics port introduced here gets its own treatment in Teaching a Bot to Walk (coming soon). For the wider system, start with A Server That Plays Itself.
Every agent system eventually invents the same vocabulary and then ruins it. Somebody says “task,” somebody else says “action,” a third person says “behavior tree node,” and within a month all three words mean all three things and nobody can review a pull request. I have watched it happen more times than I want to count.
Westworld of Warcraft has four words. They are load-bearing and they are not synonyms.
One Sentence, Four Kinds of Thing
Take one sentence a person might say about a game: “I ran Upper Blackrock Spire last night.”
Unpack it and you get four completely different kinds of thing:
The run itself - hours long, involved nine other people, had a name.
Getting to the entrance - a discrete goal with a definite end state, composed of many smaller things.
Walking through the Burning Steppes - a repeating loop with stuck detection, re-pathing, and a way to give up.
Pressing the forward key for one frame - the smallest thing a player can actually do.
They differ in duration by six orders of magnitude, they differ in who decides them, and - critically - they differ in whether anything outside the bot process needs to know about them. Flatten them into one concept and you get a system where a test cannot tell whether it is asserting on a strategy or a keystroke.
The Four Layers
Layer
Definition
Crosses the wire
Activity
A major, usually dynamic event supporting any number of characters: a raid, a battleground, a dungeon run, a multi-hour farm.
No
Objective
One high-level state change, composed of tasks. Travels as ObjectiveMessage.
Yes - the only one
Task
One behavior-tree node on a last-in-first-out stack, driving a single state change with verification and failure handling. Pushes child tasks.
No
Action
An atomic local primitive: one memory read, one bit write, one opcode send, one key press.
No
The single most useful line in the whole system is the wire column. Exactly one layer is observable from outside the bot process, and everything about testing, debugging, and service boundaries follows from that.
Where People Get It Wrong
The recurring mistake is putting compound operations in the Action layer because they feel atomic from the caller’s side.
Looks atomic
Actually is
Why
MoveToCoord(coord)
A Task
Loops position reads, movement bit writes, and heartbeat opcodes over many ticks with stuck detection
UseAbility(id, target)
A Task
Checks the global cooldown, sets the target, sends the cast opcode, then verifies the cast result
InviteToParty(player)
A Task
Sends the invite, then polls group membership until accepted or timed out
LootCorpse()
A Task
Opens the loot window, enumerates slots, takes items, verifies bag deltas
The rule that settles every argument: an Action cannot fail partway through. It either happened or it did not. If a thing can be half-done, it is a Task and it needs verification.
Convenience wrappers on the object manager - MoveToAsync, UseAbilityAsync, TurnInQuestAsync - are all Task-level. They exist because writing the same seven-Action sequence in twelve places is worse than naming it once. Naming it does not make it atomic.
Why the Wire Layer Is Deliberately Expensive
The set of objective types a state manager can request is a closed enum. Adding one costs five coordinated edits:
Add the value to the protobuf definition.
Add the mirrored value in the managed enum.
Add the mapping in the dispatcher.
Add the sequence builder for both runtimes.
Regenerate protobuf for every consumer.
That is annoying on purpose. Every time someone reaches for a new objective type, the cost forces the question: could this be a new Task that composes existing objectives instead? Almost always, yes.
The alternative - an open-ended wire vocabulary - produces a protocol that grows one verb per feature until nobody can enumerate what a bot can be asked to do. A closed set of verbs that everyone can read is a feature - twenty-six defined today, with IDs reserved through sixty-three for the ones the roadmap will want.
Here is the current shape of that vocabulary:
publicenumObjectiveType{Travel=0,// arrive at a named-location positionInteract=1,// open conversation, click an NPCAcceptQuest=2,TurnInQuest=3,Kill=4,// kill N of a creature entryCollect=5,// gather N of an itemUseGameObject=6,// chest, door, lever, herb, ore, fishing poolCastSpell=7,Escort=8,EncounterTrash=9,// a dungeon or raid trash legEncounterBoss=10,Loot=11,Queue=12,// battleground or dungeon queueCap=13,// node or flag captureHold=14,// node defenseCraft=15,Train=16,Bank=17,Mail=18,Auction=19,Vendor=20,Rebind=21,// hearthstone bindEquip=22,Loop=23,// gathering route, hotspot loop, any "until X" sweepGroupForm=24,WorldEventStage=25,// reserved through 63}
Read that list and you can predict what the server population is capable of without reading any implementation. That is the point.
From a Sentence to a Keypress
Take the dungeon run from the opening and trace it all the way down.
GoToTask is the universal child. Almost every task in the system eventually needs to be somewhere else first, so movement is factored out once. If you are writing a new task and you find yourself writing pathing, stop.
One Action, two completely different implementations. In the foreground runtime ReadPlayerPosition is a plain memory read off the client’s own object manager. The background runtime has no memory to read, so coordinates come from the movement block of SMSG_UPDATE_OBJECT and the MSG_MOVE_* traffic, and its object manager holds the latest value. Neither one reaches for an update field, because position has never been one. The task above does not know or care which implementation answered.
The last Action is a verification read. Arriving somewhere is exactly the kind of operation where a naive implementation sends the movement opcodes and then sleeps for the estimated travel time. That is a blind sequence, and blind sequences are banned. The task is not arrived until a position read puts the character inside the radius. The same shape covers the interactive objectives: a UseGameObject task sends CMSG_GAMEOBJECT_USE at a chest or a lever and then reads that gameobject’s own state byte, rather than assuming the click landed.
Only one line crossed a process boundary. The state manager said “reach Flame Crest.” It did not say how, and it never hears about the re-path around the Burning Steppes patrols.
Tasks push children; objectives do not. An objective builds exactly one head task. That task may push a whole subtree. This keeps the recursion in one place instead of spread across two layers.
Objectives Are Generated, Not Written
Here is the part that surprises people.
The activity catalog is 86 hand-authored rows - data only, no logic. A row declares what an activity is and leaves the how to the composer:
The unlock graph, which says which objectives open which other objectives.
The composer reads the world server’s own tables - quest templates and their relations, creature and gameobject templates and spawns, item templates, vendor and trainer lists, area-trigger teleports, loot templates - and synthesizes an objective list for this bot at this tick. Then it prepends precondition objectives for anything the entry requirements demand but the bot does not have.
The consequence is worth stating plainly: nobody wrote a script for Wailing Caverns. Two bots assigned the same catalog row get different objective sequences, because one already did the pre-quest and the other did not, and because one is a druid who can skip a fight the other has to take.
That is also why the world database is read-only from the bot side. Every mutation goes through the server’s own administrative interface. The composer treats the world as a fact source, and a fact source you write to is not a fact source.
Tasks, Verification, and Failure
A task is a behavior-tree node with three responsibilities: drive one state change, verify it, and fail informatively.
publicinterfaceIObjective{stringId{get;}// "ubrs.reach-flame-crest"ObjectiveTypeType{get;}IObjectiveEndStateEndState{get;}// predicate over the snapshotIReadOnlyList<ObjectiveGate>Gates{get;}// start-time preconditionsIBotTaskBuildHeadTask(BotTaskContextctx);boolCheckCompletion(WoWActivitySnapshotsnapshot);voidOnHeadTaskTerminal(BotTaskStatusterminal,string?reason);}publicinterfaceIObjectiveEndState{boolIsSatisfied(WoWActivitySnapshotsnapshot);stringDiagnosticLabel{get;}// "QuestLog[slotForQ132].Counter >= 8"}
DiagnosticLabel is the small detail that pays for itself weekly. When a bot stalls, the operator console does not say “task failed.” It says the bot was waiting on QuestLog[slotForQ132].Counter >= 8 and the counter is at 5. That is the difference between an hour of log archaeology and a ten-second answer.
The stack discipline is equally deliberate. Tasks push and pop on a last-in-first-out stack. A parent pushes GoToTask, GoToTask completes and pops, the parent resumes. There is no global “what is the bot doing” variable to get out of sync - the top of the stack is the answer, and it is published on every snapshot.
The Contract with Tests
Because objectives are the only thing on the wire, they are the only thing a test may legally drive, and snapshots are the only thing a test may legally read.
A test declares an activity, lets the composer and the resolver do their work, and asserts on published snapshot fields:
current_activity_id // "dungeon.ubrs"
current_objective_id // "ubrs.reach-flame-crest"
current_objective_type // Travel
current_task_name // top of the task stack
advice_log[] // what was suggested, and whether it was used
That last field is the advisory layer’s paper trail, and it gets its own post in Advisory, Not Authoritative (coming soon).
A test that constructs an ObjectiveMessage in its own body and dispatches it is not testing the bot. It is remote-controlling it, and it has silently skipped every layer that decides what to do - which is where the interesting regressions live. Proving a World Is Alive (coming soon) works through what that leaves you with; the point here is that the rule falls straight out of the layer model.
How a Vocabulary Rots
Nothing about this model breaks loudly. It rots, one reasonable-looking pull request at a time, and it starts with a new contributor adding an objective type because that was the shortest path.
A quest requires using a specific item on a specific corpse.
Rather than compose UseGameObject and Interact, someone adds UseItemOnCorpse to the enum.
It works. It ships.
Six months later the enum has 140 values, twelve of which are near-duplicates, and the dispatcher has a switch nobody will refactor.
The system has not gained a capability. It has gained a synonym. And because the enum is on the wire, every synonym is permanent - field numbers are never reused, so the cost is paid forever.
Situation
Correct response
A behavior does not fit an existing objective type
Write a Task that composes existing types
A task needs to be somewhere first
Push GoToTask; do not write pathing
A task cannot tell whether it succeeded
The end state is missing; write one
An objective needs to push its own children
It does not; its head task does
A test needs to force a specific action
It belongs in the action-dispatch suite, not in live validation
Four Words, Still Four Meanings
The vocabulary is holding when every objective declares an end-state predicate with a human-readable diagnostic label, no task validates state with a sleep or a counter or a fixed repeat count, the snapshot publishes the activity, the objective, and the top of the task stack on every tick, and the composer reads the world database without ever writing to it. Those I check the way you check a lock: by trying the door occasionally and moving on.
The measure I actually watch is the ratio. The objective enum has to grow more slowly than the task library, because the moment it does not, somebody has started encoding features in the wire vocabulary and the rot above has begun. The other one is diagnostic: a stalled bot has to be explainable from its snapshot alone, without attaching a debugger to anything. If I have to attach a debugger to find out what a bot is waiting on, the four words have collapsed back into one and I have written the system I opened this post complaining about.
Related Posts
The runtimes that execute all of this are in Two Ways to Wear a Character. The universal child task depends on the navigation stack in Teaching a Bot to Walk (coming soon). How the composer breaks ties is covered in Advisory, Not Authoritative (coming soon).