Remote MCP tools that run long: what your agent actually gets back
Most MCP demos show a tool that answers in a second. Real tools often don't: a scraper, a crawl or a report job can take minutes. We build web scrapers on Apify, so this is a problem we run into daily. We wanted to know what an agent actually receives when a remote MCP tool runs longer than one call, so we traced the raw JSON-RPC traffic against Apify's hosted MCP server ( mcp.apify.com , version…
Remote MCP tools often take a long time to complete, and the responses returned to agents during these long-running jobs are more complex than initially apparent. A YellowPages scraper on Apify took about 35 seconds to run, but demonstrated the various JSON-RPC responses an agent would receive.
Upon initialization, a scraper tool and four helper tools were made available. The helper tools, including get-actor-run, get-dataset-items, get-key-value-store-record, and abort-actor-run, hint at the server's expectation for the agent to return for results in long-running jobs.
The initial response arrived after roughly 30 seconds, while the job was still running. The response contained two parts: a structured JSON with the run ID, status, and storage details, and plain text informing the model of the job's status. This second part is a design choice that provides machine-readable state and plain English instructions for the model to follow.
When polling for completion, a ceiling exists – the server caps the wait time at 45 seconds. If the wait exceeds this limit, an MCP error is returned. A validation error also occurs if the wait time is set to 30 seconds, but a successful result is returned if the wait time is reduced to 30 seconds. This discrepancy emphasizes the need for clients to handle validation problems separately from protocol errors.
Results can be retrieved in pages, which is helpful for large jobs to manage the context window. In a case where a Google Maps scraper ran for 157 seconds and finished with a SUCCEEDED status but returned no items, the agent must treat SUCCEEDED as not equal to "done" and check the item count. A result with SUCCEEDED status but zero items should be considered a warning, not a success.
To handle long-running MCP tools, a checklist includes expecting an early return, storing the run ID from the first response, polling within the server's limits with an overall deadline, checking protocol errors and isError, reading results in pages to protect the context window, and treating SUCCEEDED with zero items as a warning.
If the user changes their question, abort-actor-run should be used to save the remaining job cost. Finally, limiting the number of tools available to the model can help prevent the model from selecting the wrong tool and limit its potential impact.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.