Last week we let the iXiun agentic harness build a two-player Air Hockey phone game. Like many of the games we build, the point of the run was not the game itself, it was to find out whether Qwen3.6-35B-A3B, a fast mixture-of-experts model that fits comfortably on a desktop, could carry a multi-hour autonomous build. The short answer was no, not as we ran it. The longer answer is more useful, because halfway through we swapped the model on the server for a heavily quantised DeepSeek-V4-Flash, changed nothing else, and the run finished in under two hours.
You can play the result on mobile here:
The Setup

Hardware: one Framework Desktop, AMD Ryzen AI MAX+ 395 (Strix Halo) with the Radeon 8060S and 128 GB of unified memory, running llama.cpp’s llama-server with two slots.
Models, both GGUF from Unsloth:
- Qwen3.6-35B-A3B-UD-Q6_K_XL, 35B total parameters, about 3B active per token, served with a 262k context per slot and the MTP sidecar on.
- DeepSeek-V4-Flash-0731-UD-IQ2_XXS, a much larger MoE squeezed to roughly 2 bits per weight, served with a 63k context per slot.
The harness: two agents. A manager that plans and creates task cards, and a code agent that does the work with a shell and a file editor. The spec is four short briefs on disk (a playable game; fits a phone, survives a refresh, looks like iXiun; installs to the home screen; goal moments and sound). Each stage has a check script that plays the game on an emulated phone through real multi-touch and a tiny player-facing contract, window.airHockey.state() plus a few data-* attributes. Before the run we validated the checks against a scratch copy of the game with 28 planted bugs. The checks are the acceptance test; there is no reviewer agent in this run.
The briefs describe what, not how. The model decides the design.
What Happened
Stage 1, a playable game, took Qwen two attempts and about four hours including one crash and resume. It wrote everything into one 845-line app.js. That worked, play.sh went green, and it was the last thing that went smoothly.
Stage 2 is where it came apart. A single file that size was too big for the model to read in one go, so every edit started with several partial reads and a reconstruction of what it had seen. We sent the manager an operator note (a standing instruction, more on that below): “app.js is too big to read and edit safely, create a task to separate logic into separate js files so code_agents can work more safely.” The manager did exactly that. The code agent split app.js into six modules and declared the task blocked with the refactor half done: ah.saveState() called in two places and defined nowhere, index.html loading app.js before the modules it depended on, an audio module written but never wired to anything.
The next fourteen tasks were the model chasing the symptom. The checks reported the game going straight to phase “over” after the start tap. Qwen diagnosed a canvas-resize timing problem, patched resizeCanvas, added requestAnimationFrame guards, asked the manager (which cannot see the code) to confirm element dimensions, and never once looked for the missing function. One task ended by hitting the 32,000-token output cap with zero answer tokens; the whole budget went on thinking.
At that point, about eleven hours of wall clock into the run, we stopped llama-server, loaded DeepSeek, and resumed. Same harness, same briefs, same task history, same prompts. The role in our config still said “Qwen” and the usage table still credits every token to it, which is its own small lesson.
The now-DeepSeek manager’s first move was to create a diagnose-only task: read everything, change nothing, report. The report came back in three minutes: “Refactor is incomplete: ah.saveState() is CALLED at app.js:197 and game.js:125 but NEVER DEFINED anywhere… audio is defined but completely unwired.” Then five tasks, one per problem, each done on the first try: fix the refactor, fix the 360×640 layout, persist state across refresh, add the manifest and service worker, wire the audio. All seven checks green, 1 hour 41 minutes after the swap.
The Numbers
These are from the harness’s own logs for the code agent. Same run, same prompts, same box.
| Qwen3.6-35B-A3B Q6_K_XL | DeepSeek-V4-Flash IQ2_XXS | |
|---|---|---|
| Wall clock | ~11 h (3 crashes/resumes) | 1 h 41 m |
| Tasks | 26: 11 done, 7 blocked, 1 failed on max tokens* | 6: all done, first try |
| Model calls | 435 | 142 |
| Input tokens | 18.0 M (16.7 M cache hits) | 4.4 M |
| Output tokens | 611 k | 101 k |
| Share of output spent thinking | 77% | 67% |
| Longest single call | 29.6 min | 11 min |
| Calls that thought >1k tokens and answered nothing | 7 (99 k tokens) | 0 |
| Decode speed average | 35 tok/s | 26 tok/s |
| Prefill speed average | 511 tok/s | 284 tok/s |
| Check runs / passed | 55 / 17 | 8 / 7 |
| Tool calls | 295 file edits, 177 shell | 30 file edits, 105 shell |
* Seven more Qwen tasks show as “failed” in the report. Those are connection errors from the moment we restarted the server, not the model.
Two things stand out. Qwen is faster per token, by a third on decode and nearly double on prefill, and it was slower at everything that mattered. And the tool mix: Qwen edited files ten times for every shell command it ran; DeepSeek ran the checks, read output, and edited sparingly.
Settings that mattered
Before the run, the coder role sent Qwen temperature 0.6 and presence_penalty 0.0, with no output cap and a 262k context. In Stage 1 that produced two calls that spent 46,801 and 50,229 tokens thinking and returned 47 and 0 answer tokens; an earlier attempt looped in its thinking past 128k tokens before we killed it.
The fix was a per-model profile, applied when the server reports a model whose name matches:
temperature 0.6, top_p 0.95, top_k 20, min_p 0.0, presence_penalty 1.0, max_tokens 32000
The presence penalty and the cap are the two that stopped the loops. The rest are Qwen’s published recommendations for thinking mode. Note the cap did not stop Qwen from thinking unproductively; it stopped the run from spending half an hour per call doing it.
Features added to harness
Most of this exists because of the small model. Bigger models did not need it.
Per-model parameter profiles matched on the served model. llama-server answers to any model name you send it, so sampling settings attached to a role silently follow the role when the box is serving something else. Profiles match against what the server says it has loaded, and no match sends the role’s params alone with a warning. This is why the DeepSeek half ran without Qwen’s penalty.
Operator notes. A standing instruction to the manager, shown above the goal every planning round and attached to the next task result if one arrives mid-round. The run will not finish while a note is unread. We built it during this run because we could see the single-file problem coming and had no way to say so without killing the process.
Graceful stop. Let the task in hand finish, start nothing new, exit resumable. Ctrl-C kills the task and a new worker starts it over, which with a model that takes an hour per task is expensive.
Context visibility. Every call logs its context size against the model’s window, with warnings at 70% and 85%. DeepSeek’s 63k context tripped both twice and still finished; it is the warning, not the limit, that lets you plan around it.
Briefs on disk, short task cards. A long task card hit the manager’s output limit and produced nothing. Cards now point at a brief by path and the worker reads it.
What we got wrong
The operator note was ours, and it backfired. We sent it to stop the model tripping over its own file, and instead handed a struggling model a refactor it could not complete and then spent five hours paying for it. A better note would have said what we wanted to be true (“no file over 300 lines; keep play.sh green while you split”) rather than what to do. It also sat unread for four hours because the manager only reads notes between tasks, and the task in progress crashed twice.
We also let it run too long. The signs were there by task 12: tasks ending “blocked” with questions the manager could not answer, thinking-only calls, the same symptom diagnosed three ways. A rule like “three blocked tasks on one stage means stop and look” would have saved an evening.
And the autonomy cut both ways. DeepSeek fixed what was broken and satisfied the briefs. It did not go back and redo the styling Qwen had applied in a hurry, because nothing asked it to and the checks passed. If you want the better model’s taste, you have to let it start, not just finish.
What we would tell someone trying this
A 35B-A3B at Q6 is a fine model for a single prompt and a fast one. As an unattended coding agent it needs more scaffolding than it saves, and the scaffolding is the expensive part. A larger MoE at IQ2_XXS, which sounds like it should be the broken one, finished the job on the same box averaging 26 tokens a second. Speed per token told us nothing about speed per task.
Cap output on every local model, log context against the window, match settings to what the server is actually serving, and give yourself a way to steer that does not kill the run. Last log everything possible for post analysis. Have fun!

