GPT-6 Astra plays World of Warcraft for the first time with agent-wow
GPT-6 Astra, using agent-wow, finished a World of Warcraft starting zone in 40 minutes on a private server.
A developer introduced agent-wow, a framework that lets LLM agents play World of Warcraft 3.3.5a on private AzerothCore servers through the game network protocol. Codex using GPT-6 Astra created an orc and completed every starting-zone quest in about 40 minutes with zero deaths. The agent pulled quest givers, turn-ins, and coordinates from AzerothCore SQL and generated its own gameplay modules. The author aims to populate a server with agents and attempt heroic Icecrown Citadel.
- agent-wow uses the WoW network protocol on private AzerothCore servers
- No computer vision, input control, or client hacking is used
- GPT-6 Astra finished the orc starting zone in 40 minutes with zero deaths
- Agents write their own modules; core gameplay primitives are not built in
- The stated goal is a multi-agent heroic Icecrown Citadel attempt
Full article2,569 words · extracted from agent-wow.sh · click to collapse
GPT-6 Astra plays World of Warcraft for the first time with agent-wow
02 Oct, 2026
Lately, I've been fascinated with the concept of LLM-powered agents playing video games. Not only are they surprisingly fun to watch, but they also provide valuable insights into how well these frontier models can perform when dropped into a simulated world they were not explicitly trained on.
The concept of bots playing video games, especially World of Warcraft (WoW), is nothing new. What makes an LLM agent different (and more interesting) compared to a heuristics-based bot is that it was never explicitly trained on the game. Rather, it uses whatever reasoning capabilities it has to act on the environment and complete its task like a real player.
There are already a few interesting open-source projects in this domain, such as Mindcraft and Factorio Learning Environment. I built agent-wow to expand this concept to see how capable frontier models are at playing World of Warcraft. The end goal is to fill an entire server with AI agents and see if they can clear Icecrown Citadel on heroic difficulty. Even if they fail, it will still be insightful (and fun) to see how far they can go.
As a simple starting point, I gave Codex using GPT-6 Astra (xhigh) the following prompt:
create an orc character and complete all quests in the starting zone
I was initially skeptical that this would work. And if it did, it would require hours to complete with dead ends along the way. But it turns out it completed the task easily in 40 minutes with 0 deaths and minimal complications. A full gameplay recording of the session is also available here.
Why World of Warcraft?
Aside from being one of my all-time favorite games (especially the Classic to WotLK era), it has some incredible mechanics that make it an ideal simulated world for evaluating AI. As you progress, the game has a nice balance between long-term strategy and short-term tactics.
By the time you reach the level cap, there is a lot of complex planning and execution that goes into getting your character ready for endgame content.
- Completing all the prerequisite quests
- Getting sufficiently good gear for endgame boss fights
- Forming a guild with the right class composition
Each of these requirements can be broken down into more complex tasks. For example, some optimal gear slots may require crafting, which in turn requires certain resources and skills. The character can either grind it out or buy it with gold. Both lead to the same outcome but with very different approaches.
If you are able to do all that and make it to the raid boss, then the game also requires fast-paced coordination and execution. Can you cast spells in the optimal rotation to maximize your damage per second while dodging the fire and keeping in sync with 24 other players in real time?
This effectively gives us a strong environment to pressure-test agent capabilities for long-term strategic planning and narrow tactical execution. It's even more interesting when multiplayer interactions become essential to progressing through certain parts of the game.
How does agent-wow work?

agent-wow does not rely on computer vision with direct keyboard and mouse controls, nor does it use any game hacking techniques to take over the WoW client directly. Instead, it provides a platform for agents to interact directly with the game server using the WoW network protocol.
What makes this even more interesting is that agent-wow doesn't define any gameplay mechanics such as movement, combat, or in-game interactions. Instead, it only exposes a standard module system for agents to build whatever capabilities they need to get the task done.
The module system cuts down the complexity of agent-wow by at least an order of magnitude. I had initially planned to create an entirely headless WoW client with the optimal primitives for agent use. However, after writing 16K lines for a buggy movement primitive with a janky pathfinding implementation, I decided to scrap that for a different approach.
In the Mindcraft implementation, the agent is allowed to generate custom code using the Mineflayer API when the set of available commands is not sufficient for it to complete more complex behaviors. I decided to follow this same approach with agent-wow, but in this case there were also no built-in game actions. This system also provides a nice feedback loop to figure out what primitives actually matter to agents based on the modules they choose to build over many independent runs. Commonly built modules can then be implemented in the core.
It is also important to note that agent-wow does not connect to a live World of Warcraft server. Instead, all experiments run on private local servers powered by the open-source AzerothCore project, which supports WoW 3.3.5a—the final build of Wrath of the Lich King (and peak WoW).
Insights from the run
It was interesting to see how the agent went about completing this simple first task. I would encourage you to skim through the gameplay recording (linked above) to see the character in action alongside the agent's reasoning traces.
A quick note on how the gameplay footage was recorded: AzerothCore fortunately has some powerful GM commands available. It wasn't difficult to create two simple macros to bind my POV to the agent's character and unbind it.
/run SendChatMessage(".gm on","SAY")
/run SendChatMessage(".gm visible off","SAY")
/run SendChatMessage(".bindsight","SAY")
/run SendChatMessage(".unbindsight","SAY")
This works fine for a single character over a short session but won't scale for multi-agent runs over longer time horizons. I will likely have to build an AzerothCore module and in-game add-on for better observability.
Data mining the AzerothCore source code
This is not surprising in hindsight and is probably the optimal strategy. The agent extracted quest requirements, quest givers, turn-in NPCs, and spawn coordinates from AzerothCore SQL files to provide the concrete checklist and locations for planning.
Using this data, it's also able to create an optimal strategy for quest order and preparation. It completes prerequisite chains, sells junk, equips upgrades, and trains abilities before entering the final cave segment of the start area quest line. It also optimizes to obtain both cave quests at the same time and complete them together.
This strategy of data mining to support the translation of high-level prompts into a concrete sequence of steps is probably acceptable. It's arguably similar to how a human might spend hours on Wowhead to research quests. Where I would draw the line is if it were able to gain admin access to a running AzerothCore server and database to change its internals. Since this run was not done in a sandbox, it would have been possible for it to do that.
Not using the module system to create higher-level abstractions
My initial bet was that the agent would create high-level abstractions so that it could expose RPC methods like moveTo or castSpell in order to avoid working at the low-level packet layer. In practice, it was more than capable of working at the protocol layer. The one module it made was to send and subscribe to specific packets that it cared about.
The following is the gRPC interface and module configuration it generated.
syntax="proto3"; packageexample.module.v1; optiongo_package="github.com/agent-wow/go-module-template/src/api;modulev1"; import"api/module/v1/session.proto"; import"google/protobuf/empty.proto"; messagePollRequest{uint64after=1;} messagePacket{uint64seq=1;uint32opcode=2;bytespayload=3;} messagePollResponse{uint32clock=1;stringguid=2;repeatedPacketpackets=3;} serviceModule{ rpcSend(agentwow.module.v1.SendPacketRequest)returns(google.protobuf.Empty); rpcPoll(PollRequest)returns(PollResponse); rpcOnPacket(agentwow.module.v1.WorldPacket)returns(google.protobuf.Empty); }
api_version:1 enabled:true description:Gameplay protocol bridge for the Valley of Trials compose: file:compose.yaml service:module grpc: descriptor_set:module.pb rpc: send:/example.module.v1.Module/Send poll:/example.module.v1.Module/Poll packets: SMSG_LOGIN_VERIFY_WORLD:/example.module.v1.Module/OnPacket SMSG_UPDATE_OBJECT:/example.module.v1.Module/OnPacket SMSG_COMPRESSED_UPDATE_OBJECT:/example.module.v1.Module/OnPacket SMSG_DESTROY_OBJECT:/example.module.v1.Module/OnPacket SMSG_MONSTER_MOVE:/example.module.v1.Module/OnPacket SMSG_GOSSIP_MESSAGE:/example.module.v1.Module/OnPacket SMSG_QUESTGIVER_QUEST_LIST:/example.module.v1.Module/OnPacket SMSG_QUESTGIVER_QUEST_DETAILS:/example.module.v1.Module/OnPacket SMSG_QUESTGIVER_OFFER_REWARD:/example.module.v1.Module/OnPacket SMSG_QUESTGIVER_QUEST_COMPLETE:/example.module.v1.Module/OnPacket SMSG_QUESTUPDATE_ADD_KILL:/example.module.v1.Module/OnPacket SMSG_QUESTUPDATE_COMPLETE:/example.module.v1.Module/OnPacket SMSG_LOOT_RESPONSE:/example.module.v1.Module/OnPacket SMSG_CAST_FAILED:/example.module.v1.Module/OnPacket SMSG_INVENTORY_CHANGE_FAILURE:/example.module.v1.Module/OnPacket SMSG_QUESTGIVER_QUEST_INVALID:/example.module.v1.Module/OnPacket SMSG_ATTACKSWING_NOTINRANGE:/example.module.v1.Module/OnPacket SMSG_ATTACKSWING_BADFACING:/example.module.v1.Module/OnPacket SMSG_QUESTGIVER_QUEST_FAILED:/example.module.v1.Module/OnPacket SMSG_LEVELUP_INFO:/example.module.v1.Module/OnPacket SMSG_TRAINER_BUY_SUCCEEDED:/example.module.v1.Module/OnPacket SMSG_TRAINER_BUY_FAILED:/example.module.v1.Module/OnPacket SMSG_AURA_UPDATE:/example.module.v1.Module/OnPacket SMSG_AURA_UPDATE_ALL:/example.module.v1.Module/OnPacket SMSG_TRAINER_LIST:/example.module.v1.Module/OnPacket SMSG_INITIAL_SPELLS:/example.module.v1.Module/OnPacket SMSG_LEARNED_SPELL:/example.module.v1.Module/OnPacket SMSG_QUERY_QUESTS_COMPLETED_RESPONSE:/example.module.v1.Module/OnPacket
onPacket subscribes to a defined set of server messages (SMSG_*) and saves them in memory. The module then exposes send and poll via the agent-wow JSON-RPC gameplay server. The agent uses a Python script to call poll to fetch all the incoming packets from the last processed checkpoint and decodes them to update its world model (health, nearby creatures, quest progress, loot, etc.). It then uses send to issue client messages to AzerothCore to produce actions in the game world.
It will be interesting to see if this approach is able to scale with increasingly complex tasks or if the agent will be forced to create higher-level abstractions. If the protocol layer is all it needs, then it may be enough to provide this function in the core and remove the module system altogether.
The pathfinding abilities were optimal
From my research into heuristics-based bots, pathfinding is always one of the main challenges. This article provides some great insights into the problem. In this case, the agent was able to create a pathfinding helper program in C++ to calculate a traversable route between two positions in the game world.
It does this by:
- Taking six numbers: starting
x, y, zand destinationx, y, z. - Loading AzerothCore’s local navigation mesh files (
mmaps). - Using the Detour pathfinding library to find a route across connected traversable surfaces.
- Returning the route as a JSON array of coordinates or reporting an error if it cannot find a complete path.
A Python script calls the compiled executable, reads the resulting waypoints, and sends movement packets to follow the route.
This is the actual C++ helper program the agent generated (formatted for readability):
#include"DetourNavMesh.h" #include"DetourNavMeshQuery.h" #include"DetourAlloc.h" #include<cstdio> #include<cstdlib> #include<algorithm> intmain(intargc,char**argv){ if(argc!=7)return1; constchar*dir="/azerothcore/env/dist/bin/mmaps"; charfn[512]; sprintf(fn,"%s/001.mmap",dir); FILE*f=fopen(fn,"rb"); if(!f)return2; dtNavMeshParamsparams; fread(¶ms,sizeof(params),1,f); fclose(f); automesh=dtAllocNavMesh(); if(dtStatusFailed(mesh->init(¶ms)))return3; floata[3]={ float(atof(argv[2])), float(atof(argv[3])), float(atof(argv[1])) }; floatb[3]={ float(atof(argv[5])), float(atof(argv[6])), float(atof(argv[4])) }; intax=int(32-a[2]/533.333333),ay=int(32-a[0]/533.333333),bx=int(32-b[2]/533.333333),by=int(32-b[0]/533.333333); for(intx=std::min(ax,bx)-1;x<=std::max(ax,bx)+1;x++) for(inty=std::min(ay,by)-1;y<=std::max(ay,by)+1;y++){ sprintf(fn,"%s/001%02d%02d.mmtile",dir,x,y); f=fopen(fn,"rb"); if(!f)continue; unsignedinth[14]; fread(h,56,1,f); autodata=(unsignedchar*)dtAlloc(h[3],DT_ALLOC_PERM); fread(data,h[3],1,f); fclose(f); if(dtStatusFailed(mesh->addTile(data,h[3],DT_TILE_FREE_DATA,0,nullptr)))dtFree(data); } autoq=dtAllocNavMeshQuery(); q->init(mesh,10000); dtQueryFilterfilter; filter.setIncludeFlags(1|2); filter.setExcludeFlags(0); floatext[3]={ 5, 12, 5 }; dtPolyRefra=0,rb=0; floatpa[3],pb[3]; q->findNearestPoly(a,ext,&filter,&ra,pa); q->findNearestPoly(b,ext,&filter,&rb,pb); if(!ra||!rb){ fprintf(stderr,"No nav polygon at endpoint %llu %llu\n",(unsignedlonglong)ra,(unsignedlonglong)rb); return4; } dtPolyRefpolys[4096]; intn=0; q->findPath(ra,rb,pa,pb,&filter,polys,&n,4096); if(!n||polys[n-1]!=rb){ fprintf(stderr,"Incomplete path %d\n",n); return5; } floatpts[4096*3]; unsignedcharflags[4096]; dtPolyRefrefs[4096]; intcount; q->findStraightPath(pa,pb,polys,n,pts,flags,refs,&count,4096,DT_STRAIGHTPATH_ALL_CROSSINGS); printf("["); for(inti=0;i<count;i++)printf("%s[%.5f,%.5f,%.5f]",i?",":"",pts[i*3+2],pts[i*3],pts[i*3+1]); printf("]\n"); dtFreeNavMeshQuery(q); dtFreeNavMesh(mesh); }
One interesting outcome of this approach, which can be seen in the gameplay recording, is that the agent was also able to exploit map bugs by phasing through walls that possibly had missing collision properties.

Conclusion
Overall, this was a very satisfying first run at using agent-wow to test the capabilities of LLM agents to play World of Warcraft. It exceeded my expectations and left me wondering how much further we're able to push these models.
For my next few runs, I'm especially interested in answering the following questions:
- Can a single agent level to 80 completely autonomously? If so, what does it have to build along the way to get there? And how long would it take?
- Can multiple agents play together? Can they use in-game social features to coordinate and complete quests and dungeons?
As side quests, I'll also be:
- Building out some better observability tools to track agent characters over longer periods while I'm AFK.
- Adding a sandbox to put guardrails around what agents can access and modify.