END-TO-END · SHIPPED

Hellpod Companion

A loadout recommendation tool for Helldivers 2 that explains every pick and lets you swap any of them. A solo, end-to-end, spec-driven project built for the education and experience.

Role
Product strategy · UX design · spec authorship · build
Year
2025 – 2026
Method
Spec-driven design and development. Written PRDs, decision logs, and a phased implementation plan carried the whole build.
Status
Shipped and live at hellpodcompanion.app.

00 — At a glance

One line.
A loadout recommendation tool for Helldivers 2, built solo end to end as a spec-driven side project.
What this page proves.
That the method carries. Written specs, decision logs, and a phased implementation plan produce a shipped product, and the receipts are here to read.

01 — Why This Project

A learning project, on purpose.
The AI-assisted build was the method I wanted to be good at. Reading about it wasn't going to teach me anything I hadn't already absorbed. I needed a project where I would actually sit in the seat, run the tools every day, and let the method's real shape emerge from doing the work. So I picked a target with no external stakes, no timeline pressure, and no one waiting on the outcome. The goal was the method. The product was the excuse to run it.
Why Helldivers 2 was the right domain.
A learning project only teaches you something if you can tell when the tool is wrong. Helldivers 2 is a co-op shooter with a large space of stratagems, weapons, and armor perks that interact in mechanically specific ways, and I had put enough hours in to know when a recommended loadout felt right and when it felt like nonsense. That check is something I couldn't have faked in an unfamiliar domain. The community wiki gave the AI a corpus to draw from, and the game gave me a nightly stress test: run the recommendation, drop into a mission, see whether the tool understood the thing it had just talked about. No unfamiliar user was going to get burned by a wrong answer, because the whole thing was a game.
What I thought "AI-assisted build" meant going in.
The going-in premise was that the shape would be simple. Feed the model the wiki, describe the recommender I wanted, and watch it assemble the pieces. The AI would carry the domain knowledge, and I would supervise the interface work and the copy. The tools had gotten good enough, I assumed, to do the substance of the build if you handed them a clean enough brief.
What it turned out to need.
That premise did not survive contact with the work. The game's mechanical nuance was not the kind of thing you could hand off to a model with a wiki attached and expect back in usable form. Every serious call, from how to represent a stratagem's tradeoffs to what a "supporting" pick actually did in a four-person squad, needed someone in the chair with hours in the game to hold the domain. What the method actually is, once you strip away the sales pitch, is closer to a working relationship: a junior next to a mentor. The junior moves fast and cheerfully takes on the tedious parts. The mentor keeps the plan straight, calls the domain, and rejects the plausible-looking wrong answer. Let the junior drive and you get plausible-looking wrong answers all the way down. Never let it execute and you get a beautifully specified thing that never ships. The whole practice is holding both hands on the wheel.

02 — What Shipped

The tool, in one sentence.
Hellpod Companion lives at hellpodcompanion.app. You give it your mission context, it returns a Helldivers 2 loadout with a written rationale for each pick, and if you disagree with any single element you can swap it out without regenerating the rest. Two quiet surfaces round it out, Saved for the loadouts you keep and Settings for the defaults, both intentionally out of the way.
The Recommend flow.
The core loop is a short guided intake. Four steps, each one a single question, each one answered with tappable chips drawn from the game's own vocabulary. The pattern is deliberate: the tool asks in the game's language and answers in the game's language, so the domain stays consistent from what you input to what comes back.
The Build Your Loadout wizard on step 1 of 4, with three faction options: Terminids, Automatons, Illuminate. An analytics consent banner is visible at the bottom.
Step 2 of 4 after selecting Terminids. Difficulty picker showing options 1 through 10.
FIG 02.1Two of the four intake steps. Chips are drawn from the game's vocabulary, not a generic taxonomy.
Step 3 of 4 with Terminids and Difficulty 7 selected. Planet picker showing eight options: Bore Rock, Gacrux, Gar Haren, Gatria, Nivel 43, Omicron, Pandion-XXIV, Phact Bay.
Step 4 with Bore Rock selected. An Add operation modifiers field for Mission Conditions, and a Mission Type picker listing over a dozen mission options including Blitz Search and Destroy, Cleanse Infested District, and Eradicate Terminid Swarm.
FIG 02.2The remaining two intake steps. The whole intake reads as one short conversation, not a settings screen.
The result.
The recommendation is a full loadout, presented top to bottom on one page. Each pick has a rationale sitting one tap away, so you can see why a particular grenade or support weapon was chosen for the run you described. The rationale is the load-bearing piece of the product. Without it the tool is a slot machine. With it, you can either trust the reasoning or push back on it before you drop into a mission.
Top of the Recommended Loadout page for Terminids on Bore Rock, Difficulty 7, mission launch-icbm. The armor pick (AC-1 Dutiful, tagged survivability and defensive) and the primary weapon (SMG-203 Gallant, tagged anti-armor, crowd control, mobility), each with swap and expand controls.
Bottom of the same Recommended Loadout page. The throwable (G-123 Thermite), four stratagems (Guard Dog, Eagle Cluster Bomb, Patriot Exosuit, TX-41 Sterilizer), and one booster (Hellpod Space Optimization), each with tags and swap and expand controls, plus Save Loadout, export, and Re-roll actions.
FIG 02.3A full recommended loadout, top to bottom.
The rationale drawer expanded on the primary weapon, showing the plain-language explanation: "Effective vs Terminids: crowd control. Anti-armor essential at this difficulty."
FIG 02.4One rationale, opened. Rationales are how the tool argues for itself.
Swap without starting over.
If you disagree with a single pick, you can open a swap sheet on that slot and see alternates the model would have considered next, each labeled with the tags that describe how it plays. Pick one and its full rationale surfaces alongside the loadout, same as the original pick's did. You are not forced to accept the whole recommendation or throw it out and rerun. The path came out of the playtest loop directly: same recommendation, one wrong element, and the cost of regenerating everything to fix one thing was too high.
The Swap Armor sheet open, listing three alternate armor options: AF-50 Noxious Ranger (light), AF-91 Field Chemist (medium), AF-52 Lockdown (heavy), each tagged defensive, with a More options button.
FIG 02.5The swap sheet on a single slot. Alternates are shown by tag; the rationale opens once you pick one.
Randomizer with a Safety toggle.
There is also a randomizer for when you want to hand the whole choice off. It does not ask about the mission. It just hands you a loadout. What it does ask is whether Safety is on. On, the roll stays inside loadouts the model considers viable. Off, anything goes, chaos included. Both states are named on the surface so you know which one you asked for. The toggle is small and does one honest thing: it forces the tool to declare an opinion about what "viable" means, and lets you switch that opinion off on purpose.
The Randomizer page with Safety On selected, described as "Balanced, avoids duplicate stratagem roles." A rolled loadout: TR-62 Knight armor, LAS-17 Double-Edge Sickle primary, P-69 Veto secondary, G-123 Thermite grenade.
The Randomizer page with Safety Off selected, described as "True random, anything goes." A different rolled loadout: FS-55 Devastator armor, FLAM-66 Torcher primary, P-72 Crisper secondary, G-50 Seeker grenade.
FIG 02.6Safety on (left) constrains the roll to viable rolls; Safety off (right) opens the whole space. The tool wears its opinion on the button.

03 — The Method, Plainly

Retiring "vibe-coded."
The home card for this project used to say "vibe-coded start to finish, my first shipped AI build." I'm retiring that phrase because it undersells what the method actually is, and it misleads about what the method asks of you if you want the tool to work at the end. Vibe coding, as the term is used in the world, means feeding the AI what you want in loose prose and taking what it hands back. That is not what happened here, and it is not what I want to be known for. What happened here is closer to running a working relationship with an AI that can execute quickly but does not have the judgment about your problem that you do.
Junior and mentor, plainly.
The clearest picture I have of the method is a junior and a mentor sharing the desk. The junior is fast, tireless, doesn't get bored writing boilerplate, and will happily take on the tedious parts of the job. The mentor is slower and expensive but holds the domain, keeps the plan straight, calls the hard tradeoffs, and rejects the answer that looks right but isn't. Neither one of them ships a good product alone. Let the junior drive and you get plausible-looking wrong answers, delivered confidently, on top of each other, at high speed. Never let the junior execute and you get a beautifully specified project sitting in a folder. What both of them together produce is a piece of software that is shipped, and correct on the questions that matter.
Memory files, so the junior does not start over every morning.
The AI does not remember yesterday's conversation. That is a structural fact about the tools, not a criticism. The method has to work around it. What I ran on Hellpod was a small discipline: at the end of each work session, I asked the junior to write a summary of what we had decided, what remained open, and what the current state of the code and the plan were. Those summaries lived in the repo as memory files, twelve of them across the build. Every new session started with the mentor reading the latest one, correcting anything that had drifted, and then handing it back to the junior as context. This is what let me start Tuesday's session in the same place Monday's ended, and it is what let me catch the moments where the junior remembered something wrong.
A harness for the junior's plausible claims.
The other piece was a simulation harness, a script the codebase calls sim. It runs the real recommendation engine against thousands of generated loadouts and prints the distribution of what comes out. When the junior claimed a weight change would make anti-tank picks dominant at high difficulty, or that variety across primaries had improved, the harness told me whether the claim held. It also verified that the right kinds of things surfaced at the right difficulty. The harness came with a documented reading rule for that: suppressor tags show up clearly at any difficulty, but boost tags only read cleanly at low difficulty, because the core combat tags saturate near the top of the curve. Every weight change I made through the rest of the build was validated against it before it went in.
What the method is, in a sentence.
Both hands on the wheel. The junior moves fast and cheerfully takes on the tedious parts. The mentor holds the domain, keeps the plan straight, calls the tradeoffs, and rejects the plausible-looking wrong answer. Memory files let the junior show up ready every morning. The harness turns plausible-sounding claims into settled ones in seconds. None of that is exotic and none of it is scaffolding for its own sake. It is what "AI-assisted build" turns out to need if the tool has to work at the end.
Typeset · what sat on the desk
Phase 1 problem doc · v1 PRD · v2 PRD · UX decision log · Technical decision log · Nine-phase implementation plan · PostHog instrumentation spec · Twelve session memory files · 42 commits with detailed bodies · Simulation harness · Build retrospective

04 — Where the Model Learned Something False

Three times the tool was confidently wrong.
Three times during the build the tool shipped a premise that was wrong, played correctly against that premise, and produced recommendations that looked defensible to anyone without the hours in the game to know better. Two were the AI executing correctly on a false picture of the game I had handed it. The third was the design premise itself being wrong about what the product needed to do.
Illuminate: the wrong counter, confidently coded.
The Illuminate faction fights behind energy shields, and I told the model that energy weapons were therefore the counter. That was my premise. The junior did what a junior does. It encoded the premise, weighted energy up, weighted anti-shield up, tagged six weapons anti-shield, and produced Illuminate loadouts that were internally consistent and factually wrong. Energy does not counter the Illuminate. The actual counters, once I went to the wiki to check, are explosive, which bypasses Overseer shields entirely, and anti-tank penetration, which the Harvester's shield resists explosion for and requires penetration to defeat. Neither the junior nor I had verified the premise. The junior couldn't. Game mechanics don't come out of a wiki you haven't read. I could have and didn't. The obvious fix would have been to purge everything shield-related from the faction weighting. That would have been wrong too. Three Illuminate units deal arc damage and the Lightning Spire's arc attack instant-kills you, so arc-resistance armor is a real defensive counter even though energy is not a real offensive one. The correction had to be surgical, and the mentor wrote a standing instruction to future sessions telling them not to undo the arc-resistance weighting: "This is intentional. Don't drop it."
Tag-as-proxy: a description that became a rule.
Playing the tool's recommendations in missions, I kept noticing I didn't have a reliable way to close enemy spawners. Something in the loadout might close a bug hole if the timing worked, but nothing in it was there for the job. The engine had a rule that supposedly guaranteed a closer: every loadout needed an item tagged explosive in the fourth stratagem slot. The reasoning underneath was fine. What was wrong was gating the capability on a description. Explosive-tagged sentries can't close holes. Rocket pods can't close holes. Smoke and stun grenades can't close holes. Meanwhile the Gas grenade and the Autocannon are real closers with no explosive tag on them, so the engine never considered them. And Automaton fabricators had no guarantee at all, because the rule was keyed to a tag that said "explodes," not to the capability of closing a spawner. This is where you need actual experience of the game and not just a tag on an item. The fix was a modelling change, not a tuning change. A first-class capability field got hand-curated onto forty-four items, saying which factions' spawners each item could close. The explosive tag stayed exactly where it was, still doing its scoring job. The capability gate simply stopped depending on it.
Variety: the engine did what it was designed to do, and what it was designed to do was wrong.
The third reversal is a different shape. There was no wrong premise about how the game works. The engine scored loadouts correctly and then picked a weighted random from the top ten scored candidates. That was the design. It worked. It was also a funnel. The same handful of primaries kept coming up. The same barrage stratagem kept coming up. Re-rolling, which was supposed to be the way you got a different loadout, returned near-identical loadouts. Every individual pick the engine made was defensible on the game's terms. The product was still failing at what it was for. I wrote the reframe into memory during the build: I'm not going to say that some of these picks aren't valid, but I would like to make sure that anyone using this app isn't going to get fed the same loadout each time. Validity was not the bar. The fix deleted the top-ten cap, let the candidate pool scale with the size of the category, and replaced the linear score with its square root so good items still lead but no longer dominate. All three reversals were caught the same way, by playing the game after work sessions with the recommendations the tool had produced. The harness measured the fixes; it did not find the problems. The wiki re-derivation that came after Illuminate was cleanup that leveled the AI's knowledge floor going forward, not what caught anything.
Typeset · sim run at difficulty 7, before and after the variety fix

Primaries surfaced: 10 of 50 before the fix, 25 of 50 after.

Top item share: 38 to 54 percent before, 6 to 19 percent after.

Typeset · the philosophy, recorded in memory during the build

It should NOT be a "meta machine that just spits out the same stuff everyone already is using to death," but it also should NOT be skewed by the user's own personal playstyle bias.

The goal is to surface loadouts that are balanced, effective, and make the game more fun by including things players may not typically use.

Validity is not endorsement. Cap the niche option; don't ban it.

05 — Push What You're Satisfied With

The gate.
By early June, V2 was code-complete across all its planned phases. Auth, sync, profiles, the paid-items filter, image export, account deletion. Eleven commits sat on my machine ready to deploy. One thing was not right. The transactional-email sender was still pointed at the sandbox address the provider hands you before you configure your own domain, and that address can only send email to me. Anyone who signed up with an email that wasn't mine would create an account, wait for a verification email, and never receive one. Every other flow worked. Google sign-in worked. The recommender worked. The sync worked. The paid filter worked. That one flow was broken for anyone who wasn't me.
Why I held.
The site's URL was already public. Anyone could type it in and hit whatever was live. Pushing V2 with a broken signup for non-me users was a soft-launch to organic traffic with a visibly degraded first impression, and I could not talk myself into it as a reasonable trade. The technical argument for pushing anyway was strong. Every phase had been built to be independently deployable on purpose, precisely so a partial push was possible. I rejected it. This was the first solo project of the AI-assisted arc I was starting. Pushing partially completed work is rookie dev work, and if this was how I was going to run a solo dev practice, I might as well only push the stuff I was satisfied with. Or, more simply, as I wrote it in memory during the build: "I'd rather only push something if all of it works."
What came out of it.
Both gates cleared over the next two days. The email sender got moved to the real domain. A second gate that surfaced the day after, the Supabase auth URLs still pointing at the Vercel preview domain instead of hellpodcompanion.app, got fixed the same way. All eleven commits shipped at once, followed by a manual smoke test on the live site. What sat down after was a standing rule for the rest of the arc: before a push, list every flow a real, non-owner user could touch, and confirm each one works. Anything degraded gets promoted to a pre-push gate rather than left as a launch-day footnote. It is the smallest process piece in the whole project and one I have leaned on continuously since.

06 — The Drift and the Step Away

How success quietly changed.
The project I set out to run was a learning project. The success criterion was that I ended up understanding the AI-assisted build method well enough to use it on the next thing, and I picked a domain where I could tell if the tool was wrong. That was the whole target. Somewhere after the tool shipped and the reversals were corrected, the target moved. I started thinking about how to grow the audience for hellpodcompanion.app. I put together a social media page and started planning short-form content. I looked at the PostHog dashboard more than the code. Nothing about any of that is wrong on its face. What was wrong is that none of it was the thing I had picked this project to do, and I had stopped noticing the swap.
Catching it.
The catching happened in a conversation with Claude. I do not remember which session. What I remember is describing why I felt lost about the project and hearing back a version of the sentence I needed to hear, which was that the criteria I was measuring against were not the criteria I had picked at the start. The original criteria were met. I had learned the method. I had the reversals to prove it. The tool worked. What I was actually failing at was a set of criteria I had swapped in without noticing, and that new set was about product success. That was never the point of the project. The moment I could see the swap, the lostness went with it.
Closing on purpose.
Once I could see it, the decision was easy. I closed the project. Not paused. Not shelved for a return. Closed, because the thing I picked it to teach me it had already taught me, and continuing past that point for the wrong reason would have made everything after the reversals a different kind of drift. The tool is still live. Analytics is still on. If someone finds it and uses it, I am pleased. But the project as a project ends where the learning ended, which is the honest close for a case study about a learning project.

07 — What Carried Forward

The method got named after Hellpod.
The junior-and-mentor stance was intuitive by the end of this project. The memory-file discipline was already a habit. The instinct to hold a push until every flow worked for a stranger was already an instinct. What Hellpod did not have yet was a set of names and standards for any of it, because I had been learning them by running into their absence. The next project turned the intuitions into practices. SayWhen took the method to its fullest rigor, discovery through technical handoff, each phase a real artifact with its decisions logged, and the build was scoped as a stretch and then shelved on purpose. The design was the deliverable that mattered, and it stood on its own. This portfolio site is the third pass through the same method, this time end to end and shipping. Hellpod is where the method was learned. The two projects after it are what learning it made possible.
← All Work