🕹️ agent-browser (Playwright)
The “eyes” of the design loop. Automates browser interactions — navigating, filling out forms, taking screenshots, testing web apps, and extracting data — serving as the critical link that closes the loop between generating a page and visually validate the result.
🧩 What it is / what it does
Precise definition
agent-browser is a Claude Code skill that automates browser interactions via Playwright. It exposes the tool Bash(agent-browser:*) allowing Claude to navigate URLs, take screenshots, fill out forms, test web apps, and extract data from pages—all from the terminal, without human intervention.
The skill works as a complete automation CLI. Each subcommand (open, snapshot, screenshot, fill, click…) maps to a real action in the browser. The core flow is: navigate → snapshot → interact → re-snapshot. Element references (@e1, @e2…) are extracted from the snapshot and used in the following commands.
✓ USE WHEN
- ✓ Visually validate a newly created page
- ✓ Test forms and authentication flows
- ✓ Extract data from websites without an API
- ✓ Taking screenshots for documentation or review
- ✓ Record interaction video for a demo
✗ DO NOT USE WHEN
- ✗ The task is static code analysis only
- ✗ A REST API is available for the same data
- ✗ The site blocks headless browsers (bot protection)
- ✗ You want large-scale scraping (use dedicated scrapers)
🎣 When it triggers
Description trigger
Claude invokes agent-browser when the request involves "navigate websites", "interact with web pages", "fill out forms", "take screenshots", "test web apps" or "extract information from web pages" — exactly the verbs in the skill description.
Scenarios that activate the skill
"Open the page I just created and take a screenshot" → opens localhost, captures a screenshot, returns an image.
"Fill out the contact form with test data and see if it submits" → fill + click + wait.
"Visit this page and tell me the article titles" → snapshot + get text.
"Create the page, open it, and check whether the layout looks right; fix anything that's wrong" → frontend-design generates it + agent-browser opens it + impeccable iterates.
🚀 How it improves your pages
The agent-browser is the eye closing the design loop. Without it, Claude generates HTML blindly — delivering code that may be broken in a real browser. With it, visual verification becomes part of the workflow.
📊 The closed-loop cycle with agent-browser
frontend-design / web-artifacts-builder creates the HTML page
agent-browser opens the page and takes a screenshot — Claude "sees" the actual result
impeccable iterates on what the screenshot revealed was wrong
Why it matters in design
HTML is text — visual bugs (overlapping, clipped text, wrong colors) are invisible in the code. The screenshot provides information that only exists in the rendered browser. This turns Claude from an “HTML writer” into a “designer with an eye for the result.”
✓ WITH agent-browser
- ✓ Pixel-by-pixel validated layout
- ✓ Interactions tested before delivery
- ✓ Responsive design errors detected on the spot
- ✓ Screenshot as evidence for review
✗ WITHOUT agent-browser
- ✗ Visual bugs only discovered by the user
- ✗ Forms delivered without testing
- ✗ Mobile-first never verified
- ✗ Review loop is human + slow
⚙️ How it works under the hood
Actual skill stack
Bash(agent-browser:*)
allowed-tools: agent-browser commands only
Playwright
Microsoft — cross-browser automation
--json / imagem
machine-readable or PNG/PDF
# Navegação
agent-browser open <url> # Abrir URL
agent-browser back # Voltar
agent-browser reload # Recarregar
# Snapshot (análise da página)
agent-browser snapshot # Árvore de acessibilidade completa
agent-browser snapshot -i # Apenas elementos interativos (recomendado)
agent-browser snapshot -c # Saída compacta
agent-browser snapshot -s "#main" # Escopo por seletor CSS
# Interações (usar @refs do snapshot)
agent-browser click @e1 # Clique
agent-browser fill @e2 "texto" # Limpar e digitar
agent-browser type @e2 "texto" # Digitar sem limpar
agent-browser press Enter # Teclar
agent-browser select @e1 "valor" # Dropdown
agent-browser scroll down 500 # Scroll
agent-browser drag @e1 @e2 # Drag & drop
agent-browser upload @e1 file.pdf # Upload
# Screenshots e PDF
agent-browser screenshot # Captura full page
agent-browser pdf # Exportar PDF
# Gravação de vídeo
agent-browser video start # Iniciar gravação
agent-browser video stop # Parar e salvar
# Wait / sincronização
agent-browser wait --url "**/done" # Aguardar URL
agent-browser wait --text "Sucesso" # Aguardar texto
# Sessões paralelas
agent-browser --session test1 open site-a.com
agent-browser --session test2 open site-b.com
agent-browser session list
# JSON machine-readable
agent-browser snapshot -i --json
agent-browser get text @e1 --json
# Estado salvo (autenticação)
agent-browser state save auth.json
agent-browser state load auth.json
⚠️ Warning: re-snapshot required
After any navigation or significant change to the DOM, the @refs become invalid. Run agent-browser snapshot -i again before interacting with new elements.
💬 Practical example + Ready-to-use PROMPT
Scenario: Create a landing page and validate it visually
Complete workflow: frontend-design creates the HTML → agent-browser opens it and takes a screenshot → Claude reviews the result → impeccable fixes anything that doesn't meet expectations.
Crie uma landing page para o produto "DesignBot" usando frontend-design.
Depois abra a página no browser com agent-browser, tire um screenshot
e me diga se o layout está correto — verifica se hero, CTA e footer
aparecem sem sobreposição. Se tiver algo errado, corrija com impeccable
e tire outro screenshot para confirmar.
Real example: form submission (from SKILL.md)
# 1. Abrir a página
agent-browser open https://app.exemplo.com/contato
# 2. Snapshot dos elementos interativos
agent-browser snapshot -i
# retorna: @e1 (input nome), @e2 (input email), @e3 (btn enviar)
# 3. Preencher e enviar
agent-browser fill @e1 "Maria Silva"
agent-browser fill @e2 "maria@exemplo.com"
agent-browser click @e3
# 4. Aguardar confirmação
agent-browser wait --text "Mensagem enviada"
# 5. Screenshot como evidência
agent-browser screenshot
Real example: saved authentication (from SKILL.md)
# Login uma vez
agent-browser open https://app.exemplo.com/login
agent-browser snapshot -i
agent-browser fill @e1 "usuario"
agent-browser fill @e2 "senha"
agent-browser click @e3
agent-browser wait --url "**/dashboard"
agent-browser state save auth.json
# Sessões futuras: carregar estado salvo
agent-browser state load auth.json
agent-browser open https://app.exemplo.com/dashboard
🧬 Works well with / limitations
The design loop triangle
The skill that iterates on the design with a focus on visual quality. It uses agent-browser to inspect the current state, identify issues, and fix them — forming the eye + hand of the loop.
Generates the React/HTML page with Tailwind and shadcn. After generation, agent-browser opens and validates what was created—closing the loop without leaving the terminal.
The next module. It uses Firecrawl for in-depth analysis of external sites. Combine it with agent-browser for what Firecrawl can't reach: pages that need interaction (login, SPA, lazy-load).
Other useful combinations
Generates a diagram → opens it in the browser → takes a screenshot to embed in documentation.
Generated design system → agent-browser validates responsiveness across different viewports.
animation-designer creates the CSS animations → agent-browser records the animation running in the browser.
⚠️ Real limitations
localhost:PORT).📋 Module 4.3 Summary
Bash(agent-browser:*)