Urgent.News

What's breaking now, across thousands of outlets.

Tech

Giving Browser Agents a Map of Your Product

AI coding agents can build software fast, but testing still wastes time rediscovering the product. Learn how a product surface graph makes AI QA faster.

Giving Browser Agents a Map of Your Product

Over the past year, the field of software and product engineering has undergone significant changes. Coding tools like Cursor have evolved, and Claude Code has become so proficient that most software engineering tasks now revolve around crafting effective prompts for coding agents and managing multiple agents to complete tasks more efficiently.

As a software engineer, I have found myself spending most of my time commanding these agents to generate large amounts of code. However, I soon realized that I faced a bottleneck: my agents were unable to test my features as quickly as I could run my fleet of coding agents.

To address this issue, I decided to let my agents conduct quality assurance (QA) for my product. Just as I test features manually by clicking through buttons and navigating pages, my agents were required to pass through a QA round. I used the Playwright MCP for this purpose, connected to my Claude Code open in my repository. This workflow worked well initially, but I soon discovered that my agents often spent time rediscovering how certain pages worked.

When an agent opened a page, it would take a screenshot, observe the page, then take an action such as clicking or filling out a text input. This process wasted valuable tokens as the agent treated the product as if it were seeing it for the first time. Additionally, the agent would spend a considerable amount of time understanding how certain features worked, often by searching through the codebase.

This process was inefficient, as the agent had only the CLAUDE.md file containing limited and necessary information to operate. The agent needed a digestible understanding of how the product worked and how to test it.

One engineer behind Grok Bot, Lauren Tan, has spoken about this very issue in her guide to agentic software engineering. She emphasizes the importance of giving agents all the necessary tools and resources they need to run and test things. This includes providing agents with reliable computers and development environments to test in, as well as the proper tools like Playwright MCP and a comprehensive understanding of the product's features.

However, even with these resources, agents still struggle to navigate complex products, especially when the product's context is not shared with the agent.

To overcome this challenge, Lauren introduced Feature Maps in her article. These maps are essentially markdown files that detail how certain features work within the app. By incorporating these feature maps into a /verification-skill, agents receive the detailed knowledge base of product features as context, significantly reducing the time spent rediscovering product features.

This approach improved my coding agents' testing time for features involving multiple screens from 25 minutes to just 15 minutes. However, I still encountered limitations: maintaining a full feature map required rescanning through all changes up to a certain point, and the feature maps treated features as a flat list, making it difficult to plan nested overlays and cross-surface journeys.

To address these limitations, I decided to build a knowledge graph that captures how our product works. This knowledge graph includes information about where things live, how different users navigate the platform, and the workflows they use to accomplish tasks. By giving agents a map of the product, I aimed to optimize computer-use agents for software development.

This idea is supported by research in optimizing computer-use agents, such as PG-Agent, which treats GUIs as graphs instead of linear click chains, and ActionEngine, which builds an offline state-machine of page/window states and allows agents to plan programs against that graph. These papers have shown significant improvements in task success rates, with PG-Agent achieving 95% success on a subset benchmark compared to the 66% success rate of the baseline.

Written by urgent.news from HackerNoon's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at hackernoon.com →

More in Tech

A mistake changed my career — production went down at 2am.

The worst bug of my career didn't wake me with an error. It woke me with a phone call. 2am. A paying customer, locked out of their own account, more confused than angry — which was worse.

  • Developer's system success message was not always trustworthy
  • Bug caused customers to be locked out during brief window
  • Developer learned importance of separating code writer from verifier

More from Monday 28 September →