PTENES
MODULE 4.3 Practical

🕹️ agent-browser (Playwright)

The “eyes” of the design loop. Automates browser interactions — navigating, filling out forms, taking screenshots, testing web apps, and extracting data — serving as the critical link that closes the loop between generating a page and visually validate the result.

6
Topics
25
Minutes
Practical
Level
Support
Category
localhost:3000/design [ component ] agent-browser Playwright
1

🧩 What it is / what it does

🕹️

Precise definition

agent-browser is a Claude Code skill that automates browser interactions via Playwright. It exposes the tool Bash(agent-browser:*) allowing Claude to navigate URLs, take screenshots, fill out forms, test web apps, and extract data from pages—all from the terminal, without human intervention.

The skill works as a complete automation CLI. Each subcommand (open, snapshot, screenshot, fill, click…) maps to a real action in the browser. The core flow is: navigate → snapshot → interact → re-snapshot. Element references (@e1, @e2…) are extracted from the snapshot and used in the following commands.

✓ USE WHEN

  • ✓ Visually validate a newly created page
  • ✓ Test forms and authentication flows
  • ✓ Extract data from websites without an API
  • ✓ Taking screenshots for documentation or review
  • ✓ Record interaction video for a demo

✗ DO NOT USE WHEN

  • ✗ The task is static code analysis only
  • ✗ A REST API is available for the same data
  • ✗ The site blocks headless browsers (bot protection)
  • ✗ You want large-scale scraping (use dedicated scrapers)
Playwright
Automation engine
Snapshot
Accessibility tree
@refs
Element refs
Headless
No visual interface
2

🎣 When it triggers

Description trigger

Claude invokes agent-browser when the request involves "navigate websites", "interact with web pages", "fill out forms", "take screenshots", "test web apps" or "extract information from web pages" — exactly the verbs in the skill description.

Scenarios that activate the skill

Post-build visual validation

"Open the page I just created and take a screenshot" → opens localhost, captures a screenshot, returns an image.

Form testing

"Fill out the contact form with test data and see if it submits" → fill + click + wait.

Data extraction

"Visit this page and tell me the article titles" → snapshot + get text.

Iterative design loop (key!)

"Create the page, open it, and check whether the layout looks right; fix anything that's wrong" → frontend-design generates it + agent-browser opens it + impeccable iterates.

Navigate
open URL
Fill in
fill / type
Test
check state
Capture
screenshot
3

🚀 How it improves your pages

The agent-browser is the eye closing the design loop. Without it, Claude generates HTML blindly — delivering code that may be broken in a real browser. With it, visual verification becomes part of the workflow.

📊 The closed-loop cycle with agent-browser

1. Generate

frontend-design / web-artifacts-builder creates the HTML page

2. View

agent-browser opens the page and takes a screenshot — Claude "sees" the actual result

3. Fix

impeccable iterates on what the screenshot revealed was wrong

Why it matters in design

HTML is text — visual bugs (overlapping, clipped text, wrong colors) are invisible in the code. The screenshot provides information that only exists in the rendered browser. This turns Claude from an “HTML writer” into a “designer with an eye for the result.”

✓ WITH agent-browser

  • ✓ Pixel-by-pixel validated layout
  • ✓ Interactions tested before delivery
  • ✓ Responsive design errors detected on the spot
  • ✓ Screenshot as evidence for review

✗ WITHOUT agent-browser

  • ✗ Visual bugs only discovered by the user
  • ✗ Forms delivered without testing
  • ✗ Mobile-first never verified
  • ✗ Review loop is human + slow
Seamless loop
generate → view → fix
Visual ground truth
screenshot = reality
Automated validation
without opening a browser manually
Evidence
screenshot as proof
4

⚙️ How it works under the hood

Actual skill stack

Exposed tool
Bash(agent-browser:*)

allowed-tools: agent-browser commands only

Engine
Playwright

Microsoft — cross-browser automation

Output
--json / imagem

machine-readable or PNG/PDF

Real commands from SKILL.md agent-browser CLI
# Navegação
agent-browser open <url>          # Abrir URL
agent-browser back                  # Voltar
agent-browser reload                # Recarregar

# Snapshot (análise da página)
agent-browser snapshot              # Árvore de acessibilidade completa
agent-browser snapshot -i           # Apenas elementos interativos (recomendado)
agent-browser snapshot -c           # Saída compacta
agent-browser snapshot -s "#main"  # Escopo por seletor CSS

# Interações (usar @refs do snapshot)
agent-browser click @e1             # Clique
agent-browser fill @e2 "texto"     # Limpar e digitar
agent-browser type @e2 "texto"     # Digitar sem limpar
agent-browser press Enter           # Teclar
agent-browser select @e1 "valor"  # Dropdown
agent-browser scroll down 500      # Scroll
agent-browser drag @e1 @e2        # Drag & drop
agent-browser upload @e1 file.pdf # Upload

# Screenshots e PDF
agent-browser screenshot            # Captura full page
agent-browser pdf                   # Exportar PDF

# Gravação de vídeo
agent-browser video start           # Iniciar gravação
agent-browser video stop            # Parar e salvar

# Wait / sincronização
agent-browser wait --url "**/done"  # Aguardar URL
agent-browser wait --text "Sucesso" # Aguardar texto

# Sessões paralelas
agent-browser --session test1 open site-a.com
agent-browser --session test2 open site-b.com
agent-browser session list

# JSON machine-readable
agent-browser snapshot -i --json
agent-browser get text @e1 --json

# Estado salvo (autenticação)
agent-browser state save auth.json
agent-browser state load auth.json

⚠️ Warning: re-snapshot required

After any navigation or significant change to the DOM, the @refs become invalid. Run agent-browser snapshot -i again before interacting with new elements.

allowed-tools
Bash(agent-browser:*)
Sessions
Parallel browsers
State save/load
Persisted authentication
Video recording
Interaction recording
5

💬 Practical example + Ready-to-use PROMPT

Scenario: Create a landing page and validate it visually

Complete workflow: frontend-design creates the HTML → agent-browser opens it and takes a screenshot → Claude reviews the result → impeccable fixes anything that doesn't meet expectations.

READY-TO-USE PROMPT — paste into Claude Code
Crie uma landing page para o produto "DesignBot" usando frontend-design.
Depois abra a página no browser com agent-browser, tire um screenshot
e me diga se o layout está correto — verifica se hero, CTA e footer
aparecem sem sobreposição. Se tiver algo errado, corrija com impeccable
e tire outro screenshot para confirmar.

Real example: form submission (from SKILL.md)

bash — agent-browser form submission
# 1. Abrir a página
agent-browser open https://app.exemplo.com/contato

# 2. Snapshot dos elementos interativos
agent-browser snapshot -i
# retorna: @e1 (input nome), @e2 (input email), @e3 (btn enviar)

# 3. Preencher e enviar
agent-browser fill @e1 "Maria Silva"
agent-browser fill @e2 "maria@exemplo.com"
agent-browser click @e3

# 4. Aguardar confirmação
agent-browser wait --text "Mensagem enviada"

# 5. Screenshot como evidência
agent-browser screenshot

Real example: saved authentication (from SKILL.md)

bash — state save / load
# Login uma vez
agent-browser open https://app.exemplo.com/login
agent-browser snapshot -i
agent-browser fill @e1 "usuario"
agent-browser fill @e2 "senha"
agent-browser click @e3
agent-browser wait --url "**/dashboard"
agent-browser state save auth.json

# Sessões futuras: carregar estado salvo
agent-browser state load auth.json
agent-browser open https://app.exemplo.com/dashboard
Loop prompt
create → view → fix
Form test
fill + click + wait
State persistence
reusable auth.json
Screenshot proof
visual evidence
6

🧬 Works well with / limitations

The design loop triangle

✨ impeccable
Track 1 — Build

The skill that iterates on the design with a focus on visual quality. It uses agent-browser to inspect the current state, identify issues, and fix them — forming the eye + hand of the loop.

impeccable → agent-browser → impeccable
⚛️ frontend-design
Track 1 — Build

Generates the React/HTML page with Tailwind and shadcn. After generation, agent-browser opens and validates what was created—closing the loop without leaving the terminal.

frontend-design → agent-browser
🔭 website-intelligence
Track 4 — Support

The next module. It uses Firecrawl for in-depth analysis of external sites. Combine it with agent-browser for what Firecrawl can't reach: pages that need interaction (login, SPA, lazy-load).

agent-browser + website-intelligence

Other useful combinations

📊
beautiful-mermaid + agent-browser

Generates a diagram → opens it in the browser → takes a screenshot to embed in documentation.

🎯
design-designer + agent-browser

Generated design system → agent-browser validates responsiveness across different viewports.

🎞️
agent-browser video + animation-designer

animation-designer creates the CSS animations → agent-browser records the animation running in the browser.

⚠️ Real limitations

✗
Bot protection (Cloudflare, hCaptcha) — many sites block headless Playwright. In those cases, website-intelligence with Firecrawl is better.
✗
Bulk scraping — agent-browser is interactive, not for bulk tasks. To collect hundreds of pages, use dedicated tools.
✗
Source code analysis — if you want to understand a site’s HTML/CSS, Read + Grep are more efficient than snapshot.
✗
No server running — to validate a local page, the dev server needs to be running (localhost:PORT).
impeccable
loop handle
frontend-design
generates the page
website-intelligence
analysis of external websites
Bot limits
isn't invincible

📋 Module 4.3 Summary

✓ agent-browser automates Playwright via Bash(agent-browser:*)
✓ Core workflow: open → snapshot -i → use @refs → interact → re-snapshot
✓ It's the design loop's "eye": generate → view → fix
✓ Triggers when a request involves browsing, filling out, testing, taking screenshots, or extracting data
✓ Supports parallel sessions, state save/load, JSON output, and video recording
✓ Works with impeccable (iteration) and website-intelligence (in-depth analysis)
Next module
🔭
4.4 — website-intelligence
In-depth website analysis with Firecrawl—insights into the design, structure, and content of any website.