Claude Opus 5.5 built a playable Megabonk clone in one 20-hour run
Opus 5.5 can code autonomously for hours. Here's what that means for building games, apps, and long projects with Claude Code.
Picture handing an AI a single prompt, walking away for almost a full day, and coming back to a playable 3D game with character tiers, a working pause menu, and sound design that actually matches the game it's imitating. That's not a hypothetical. It's what happens when you point Claude Opus 5.5 at a long, unsupervised coding task and let it build, test, and rebuild on its own loop for close to 20 hours straight.
The target in this case was a clone of Megabonk, the breakout 3D roguelike survival game. Megabonk is a roguelike survival game where you fight your way through hordes of enemies and bosses in randomly generated maps, grabbing loot, leveling up your character, and upgrading your weapons to survive. It's become a Steam hit for good reason — it's a 3D Vampire Survivors–like that features procedurally-generated maps, automatic combat against hordes of enemies, and character progression through randomized upgrades. That's a genuinely complex game to recreate: multiple unlockable characters, tiered rarity systems, UI menus, audio, and combat balancing all have to work together.
Letting a model run unsupervised for 20 hours on something that complicated used to sound like a recipe for a pile of broken code. With Opus 5.5, it produced something close to a real game. Here's what's actually going on under the hood, and how you can try long-running builds like this yourself.
What is Claude Opus 5.5 and why does it matter for coding?
Claude Opus 5.5 is Anthropic's newest flagship model, and it was built specifically with long, unsupervised work in mind. Claude Opus 5.5 is Anthropic's strongest Opus model yet, powering long-running, highly capable agents while delivering improvements in coding and professional work.
What makes it different from previous Opus releases isn't just raw intelligence — it's efficiency over long stretches of work. Opus 5.5 is Anthropic's strongest Opus model for agentic coding, handling long-running work in large codebases including building features, debugging, refactoring, and code review. It finds the root cause before changing anything, checks its work as it goes, and explains its changes in plain language. That self-checking loop — build, test, evaluate, repeat — is exactly the pattern that let it keep iterating on a game for nearly 20 hours without going off the rails.
The benchmarks back this up. On release, Anthropic reported 66.4% on Terminal-Bench 4.0 for Opus 5.5, compared with 55.8% for Fable 5.1 and 52.3% for Opus 5. On FrontierCode v1.1 Main, Opus 5.5 reached 54.4%, versus 50.3% for Fable 5.1 and 48.0% for Opus 5. And on CursorBench 4.0, it hit 57.8%, compared to 51.8% and 46.6%. Those numbers matter because Terminal-Bench and CursorBench specifically measure how well a model handles real, messy coding environments rather than isolated quiz-style questions.
When did Opus 5.5 actually come out?
Anthropic released Claude Opus 5.5 on September 22, 2026, the first model in a new Claude 5.5 family. That timing matters if you're trying to figure out whether you're on the latest version — anything dated before late September 2026 is running the older Opus 5.
How much does Opus 5.5 actually cost?
This is where things get genuinely interesting. Frontier-level coding models have historically been expensive to run for long sessions, but Opus 5.5 flips that script.
Claude Opus 5.5 is built for long-running agentic coding and knowledge work, priced at $4 / $20 USD per million input / output tokens. That's a meaningful drop from the previous generation — down from Opus 5's $5 and $25 per million tokens. Anthropic frames the savings even more aggressively for typical workloads: pricing for Opus 5.5 costs an estimated 40% less to run than Opus 5 for typical workloads, while performing at the level of Claude Fable 5.1 on most work.
If you're running it through the chat interface rather than the API, access starts with a standard subscription. Opus 5.5 is available on Claude for Pro, Max, Team, and Enterprise users — meaning the $20/month Pro plan is enough to get hands-on with the same model that ran that 20-hour build.
For developers calling the API directly, the model ID is straightforward to reference: the API model ID is claude-opus-5-5, with a 1M token context window, max output of 128K tokens (300K on the Batch API, in beta), and a reliable knowledge cutoff of June 2026.
How do you actually run a 20-hour coding session?
A single chat message isn't going to run for 20 hours on its own — you need the right surface for that kind of sustained, autonomous work. That's where Claude Code comes in, Anthropic's dedicated coding agent that can operate independently for extended periods.
Claude Code builds the plan, asks clarifying questions, and handles work that runs for hours or days, with you setting the direction as the architect and orchestrator while Claude Code does the work. That's the structural piece that makes a 20-hour game-building marathon possible: you're not babysitting every line of code, you're setting an objective and letting the agent iterate.
What's the simplest way to get started with Claude Code?
You've got a few install options depending on your platform. The recommended method, straight from Anthropic's own repository, is a one-line install:
curl -fsSL https://claude.ai/install.sh | bash
On Windows, you'd run the PowerShell equivalent instead:
irm https://claude.ai/install.ps1 | iex
Once installed, you cd into a project directory and type claude to start a session, authenticating either with a membership or by paying per call using an API key. From there, the official quickstart guide walks through your first real task step by step.
How do you kick off a task that runs for hours without you watching it?
This is the part that makes marathon builds like the Megabonk clone possible: Claude Code doesn't have to live only in your terminal. You can run Claude Code in your browser with no local setup, kicking off long-running tasks and checking back when they're done, working on repos you don't have locally, or running multiple tasks in parallel.
That cloud-based session persists independently of your laptop. A session keeps running even if you close the browser or shut your laptop — which is exactly how you'd let a build run overnight (or for 20 hours straight) without needing your machine on the whole time. You can check in on progress from the Claude Code web interface or pull a running session back into your terminal later.
What does this mean for building your own passion project?
The headline takeaway isn't "AI can clone a specific indie game." It's that the loop of build-test-debug-rebuild, which used to require constant human steering, can now run unattended for long stretches and still produce coherent, playable output. That opens up a different way of approaching side projects:
- Write a detailed spec first. The more specific your prompt about mechanics, art style, and systems (character tiers, upgrade rarities, UI elements), the less the model has to guess during those unattended hours.
- Let it loop on testing. Ask explicitly for a build-then-test cycle so the agent catches its own bugs before you ever open the project.
- Check in periodically, don't hover. Long-running sessions are designed for you to walk away. Use the web interface or scheduled check-ins rather than watching every token stream by.
- Start smaller before going 20 hours. Run a one or two hour session first to see how the model handles your specific codebase or game engine before committing to an overnight run.
If you want to compare how Opus 5.5 stacks up against other models for your own workflow, the Anthropic pricing documentation breaks down costs across the full lineup, including cache pricing that can cut repeated-context costs significantly on long sessions.
The real story here isn't the game
A solo developer's roguelike getting cloned by an AI agent is a fun headline. But the more useful insight is what it reveals about where coding agents are headed: longer unsupervised runs, cheaper tokens, and fewer babysitting requirements. The gap between "AI that helps you code" and "AI that codes while you do something else entirely" keeps shrinking — and the next long-running build worth trying might be the one sitting in your own backlog.