How AI agents actually work (and how to build one that won't wreck your files)
Learn what AI agents really are, how MCP connects them to your tools, and the exact permission habits that keep an agent from going rogue.
Picture a folder full of messy screenshots, half-written notes, and three different versions of the same spreadsheet. Now picture typing one sentence and watching that chaos turn into a finished slide deck — no copy-pasting, no manually opening each file, no formatting by hand. That's not a demo reel fantasy anymore. It's what a properly set-up AI agent does on a Tuesday afternoon.
The gap between "AI that answers questions" and "AI that does the work" is closing fast, and most of that gap closes through one unglamorous piece of plumbing: how the agent connects to the tools, files, and apps you already use. Get that part right, and an agent turning a folder into a presentation stops being magic and starts being a Tuesday.
This guide breaks down what an AI agent actually is, how it talks to your files and browser, and — most importantly — the permission habits that keep you in control of it instead of the other way around.
What actually makes something an "AI agent" and not just a chatbot?
A chatbot answers. An agent acts. That's the whole distinction, but it has big implications.
A regular chat model takes your prompt, generates text, and stops. An agent takes your instruction, breaks it into steps, calls tools to execute those steps (open a file, run a search, click a button, write a new document), looks at the result, and decides what to do next — in a loop, without you re-prompting it after every action. The "thinking" and the "doing" are stitched together into one continuous process instead of a back-and-forth you have to manage manually.
What makes this possible under the hood is a standard way for the model to discover what tools are even available to it. That's where the Model Context Protocol comes in.
How does an agent know what it's allowed to touch?
This is the part most beginners skip past, and it's the most important concept in the whole course of modern agent-building. Model Context Protocol (MCP) is an open standard, introduced by Anthropic, that solves the "how does the model know what tools exist" problem.
The Model Context Protocol includes an MCP Specification that outlines the implementation requirements for clients and servers, along with SDKs for different programming languages that implement it. In practice, that means your database can expose a set of queries, your file system can expose read/write/move actions, and your browser can expose clicks and page reads — and the agent just asks each one, "what can you do?" The server answers with a list of actions, and the agent picks from that list. It never needs to know the internal details of how your database or browser actually works; it only needs the menu of actions MCP exposes.
MCP tools can expose data from the model context and perform actions with the credentials you provide, so the official guidance is to connect only to servers you trust, use least-privilege credentials, and require approval for sensitive operations. That last part — approval for sensitive operations — is the seed of everything discussed later in this guide about permissions.
If you want to go straight to the source, Anthropic's MCP documentation and the official MCP GitHub organization are the best starting points. The official MCP documentation is built using Mintlify and available at modelcontextprotocol.io.
How does an agent control a browser without just guessing?
For years, "AI automation" for web tasks meant screenshotting a page and asking a model to guess coordinates. That approach was brittle and slow. The newer generation of agent tooling actually drives the browser.
The browser use tool lets Claude navigate, read, and interact with webpages in a browser that your application runs, working with the page both through its structure — the accessibility tree, elements, forms, and tabs — and through screenshots and viewport coordinates. That distinction matters: instead of blindly clicking pixel coordinates, the agent can read the actual DOM structure of a page, find a specific button by its role, and click it reliably even if the layout shifts slightly.
This is a real step up from the old "automation tool plus screenshot" workaround. This kind of browser automation differs from basic computer use because it's specifically optimized for web automation with DOM-aware features like element targeting, page reading, and form manipulation.
If you're using Claude specifically, the setup is more plug-and-play than you'd expect. Claude Code integrates with the Claude in Chrome browser extension to give you browser automation capabilities from the CLI or the VS Code extension, letting you build your code, then test and debug in the browser without switching contexts. One detail worth knowing before you hand over the keys: Claude opens new tabs for browser tasks and shares your browser's login state, so it can access any site you're already signed into.
That's powerful — and it's exactly why permissions matter so much, which is covered below.
What's the difference between browser use and full computer use?
Anthropic splits these into two separate tools, and knowing which one you need saves a lot of setup headaches. For tasks that stay inside webpages, the browser use tool is the closer fit because its member tools read and act on the page itself, and it doesn't need a full desktop environment.
Computer use is the bigger hammer — full screen control, mouse, keyboard, the works — for when a task lives outside the browser entirely, like a native desktop app. Claude's logic for picking a method actually follows a hierarchy: if an MCP server exists for the service, it uses that; if the task is a shell command, it uses Bash; if the task is browser work with Chrome set up, it uses that; and only if none of those apply does it fall back to full computer use.
That hierarchy is worth internalizing even if you never read another line of documentation: agents should always reach for the most specific, most constrained tool available, not the broadest one.
How do you get started with a real agent setup today?
You don't need to build any of this from scratch. A few paths exist depending on how much control versus convenience you want.
For Claude users, the fastest on-ramp is Claude Code, Anthropic's CLI-based agent environment. Getting browser control running takes just a few steps based on the official setup guide:
- Install the Claude in Chrome extension from the Chrome Web Store (verify the publisher is Anthropic).
- Launch Claude Code with the Chrome flag:
claude --chrome, or run/chromeinside an existing session to check connection status. - The integration is working when the status panel shows "Status: Enabled" and "Extension: Installed."
- If multiple browsers are connected, use
/chrome→ Select browser… to choose which one Claude drives.
For desktop-level control beyond the browser, computer use is a separate toggle: computer use is available as a built-in MCP server called computer-use, off by default until you enable it — open the MCP menu by running /mcp in a session, then find and enable computer-use in the server list.
For anyone who wants a local, subscription-free agent — especially useful if you're handling files you can't send to a third-party server — Open Interpreter is the most mature open-source option. It runs entirely on your own machine and connects to local models instead of a metered API.
Getting it running is genuinely three commands:
pip install open-interpreter
interpreter --local
Open Interpreter can be run fully locally, though you'll need to install software to run local LLMs — it supports multiple local model providers such as Ollama, Llamafile, Jan, and LM Studio. The --local flag opens an interactive menu that walks you through picking a provider. A Local Explorer was created to simplify using it locally — running interpreter --local lets you select your chosen local model provider from a list of options.
The trade-off is honest and worth stating plainly: local setup costs you time instead of money. No subscription meter running in the background, but you're responsible for hardware, model downloads, and tuning.
What permission habits actually keep an agent in check?
This is the part that separates people who use agents safely from people who end up explaining to their team why 400 files got renamed. Permissions aren't paranoia — they're how you remove decisions you don't want the agent making for you.
A few habits worth adopting from day one:
- Scope every task to a folder. Before handing off a real instruction, specify the exact directory the agent can touch and explicitly say it shouldn't go outside it.
- Say "write new files instead of overwriting originals." This single instruction prevents the most common and most painful mistake — an agent "fixing" a file by replacing it, with no way back.
- Block installs and network actions you didn't ask for. An agent that decides to install a package or send data somewhere mid-task is making decisions you didn't authorize.
- Review the tool's permission model before the first real run. Claude's level of control varies by app category — browsers and trading platforms are view-only, terminals and IDEs are click-only, and everything else gets full control. Knowing that hierarchy ahead of time tells you exactly where to expect friction.
- Watch for interrupt points. When Claude encounters a login page or CAPTCHA, it pauses and asks you to handle it manually — which is a good sign the guardrails are working, not a bug to route around.
None of this is about distrust. It's about setting the boundaries once, clearly, so you're not re-litigating them every time you run a task.
Which tools should be on your shortlist?
If you're building out an agent workflow this month, here's a practical starting lineup:
- Claude Code — best for developers who want CLI-based agent control with native browser and computer-use toggles.
- Model Context Protocol servers — the connective layer for linking any agent to databases, file systems, or custom internal tools.
- Open Interpreter — best for privacy-sensitive work or anyone who wants a permanent, subscription-free local agent.
- Claude in Chrome — purpose-built browser automation that reads page structure instead of guessing from screenshots.
Where does this go from here?
Agents are only going to get more embedded into daily tools, not less. The folder-to-presentation demo isn't the ceiling — it's the floor. The real skill to build isn't prompting harder. It's deciding, deliberately, what an agent is and isn't allowed to touch before you ever press enter on that first instruction. Start small, scope tightly, and let the agent earn more permission as it proves it deserves it.