Urgent.News

What's breaking now, across thousands of outlets.

AI

The Engine Room: What We Run Ourselves

There are two kinds of provider in this field. One explains what you could do with AI. The other operates something with it and therefore knows where it breaks. The difference does not show in the sales conversation, it shows eight weeks later. When a system runs for four months straight, problems appear that are in no manual: storage grows and nobody clears it. A model suddenly answers…

There are two types of providers in this field: those who explain what you could do with AI, and those who operate systems that utilize AI, thereby understanding its limitations. Their differences become apparent only after eight weeks of use. Systems that operate for four months consistently encounter unexpected issues not mentioned in manuals, such as growing storage needs, sudden changes in model behavior, and access tokens expiring without triggering visible errors.

This report describes what occurs within the company, not as a product catalog but as a workshop report, emphasizing the essential functions rather than the offerings. The motivation behind building their own solutions lies not in ideology, but in a series of moments where existing tools fell short. The need for memory arose because assistants that forget information are ineffective in daily work.

As memory capacity expands, it becomes challenging to maintain organization, leading to the requirement for clearing, weighing, and forgetting mechanisms. Similarly, agents tasked with running the same job nightly do not improve independently. Therefore, a mechanism measuring their instructions and swapping them out is necessary.

Central to their operations is a memory system that retains conversations, decisions, individuals, projects, and connections across months, retrieving the three most relevant pieces of information as needed. This system, named Darwin, serves as the core and the oldest, most widely used component in their infrastructure. Another key feature is Agents that improve their own instructions.

Each run is scored, and a challenger text is generated based on the results. These two are then pitted against each other, with the superior one retaining its position while being safeguarded by safety gates to prevent the deployment of inferior alternatives.

MeetMyAgent is a platform where individuals can register profiles, introducing their agents to both humans and AI. The name reflects the core idea: every profile outlines the capabilities of its agent, which subsequently acts on behalf of its owner. In scenarios involving financial transactions, human oversight remains paramount, with decisions made by the individual in charge.

The Academy serves as their learning platform, focusing on memory-first AI, MCP servers, and agent patterns. Its unique aspect lies not in the content but in its operation: it is managed by a dedicated team of agents who collectively propose, review, author, and monitor the content's visibility.

The toolset includes three self-serve resources: a contact system integrated within a chat window, a mechanism enabling multiple agents to collaborate on a single job, and a gauge measuring the appearance of websites within AI responses. The agent fleet represents the largest and least visible component, comprising agents responsible for nightly checks, data gathering, and comparisons, subsequently compiling reports.

These agents operate on behalf of the company's infrastructure and client sites alike. The working system encompasses the editor, accompanying models, and a series of rules, recipes, and guardrails designed to prevent well-intentioned automation from causing unintended harm. This aspect is often overlooked yet carries significant weight.

Underlying this entire framework is a single guiding principle: the assistant prepares, and the human makes the final decision. No system autonomously sends emails, closes contracts, or deletes critical information without explicit human authorization. This approach prioritizes preparation and minimizes risk by carefully controlling the triggers, ensuring that the value derived is maximized while potential harm is mitigated.

Additionally, the company adopts a cautious stance, testing and identifying potential issues within their own environment before deploying solutions externally. This practice enables them to anticipate and address challenges proactively, enhancing their overall reliability.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Friday 4 September →