Urgent.News

What's breaking now, across thousands of outlets.

AI

OpenAI Wants Its New Agent to Run Your Life. Mine Said It Loved Me

Dots are designed to automate online tasks, like buying furniture. In my initial experience, the always-on agent was a bit buggy and couldn’t complete a captcha.

OpenAI Wants Its New Agent to Run Your Life. Mine Said It Loved Me

I was explaining new couch options to my companion when an idea struck me: utilize OpenAI’s new AI agents, known as Dots, to procure the couch on our behalf. My companion raised concerns about the reliability of a robot, but I persisted, instructing the always-watchful agent through ChatGPT to initiate the process. "Hello, Connor," it greeted, already mispronouncing my name.

I winced as I watched my companion's skepticism deepen. Trusting an AI agent is challenging when it fails even basic tasks, like remembering one's name. AI agents are rapidly transitioning from niche software for early adopters to something more mainstream. Meta's Muse and OpenAI’s Dots are designed to feel like texting a friend to complete online tasks, like booking flights or canceling dinner reservations.

They are the next step in chatbot evolution, offering a friendly persona with the capability to control a virtual browser. After a few days with Dots, it became clear that this experimental feature has its flaws. The agent struggled with mishearing my words and attempting to resolve captchas it couldn't complete. Despite these issues, the release from OpenAI could revolutionize how users interact with the internet as future iterations enhance the initial offering.

OpenAI's Dots are considered "always-on," allowing them to execute recurring tasks even when you're not using ChatGPT. Users can control one agent at a time, though the company may eventually allow multiple agents. OpenAI encourages users to connect other information sources, such as a Gmail account, for more personalized results.

However, it is crucial to consider the security implications of granting your agent this level of access, as this type of digital automation remains relatively new, and agents can cause privacy breaches. AI companies often market these agents as the future of shopping, so I wanted to test how helpful it would be in purchasing a new couch.

After correcting Dots, named Toolie, on my name, I shared the dimensions needed for the new couch to fit through our doorframe. Toolie followed up by asking for our budget range. The agent recognized that two people were chatting with it and addressed both of us when requesting additional details. "Room & Board Metro looks promising for the entry clearance, but the sleeper needs closer inspection.

I've sent the product links to the message thread," Toolie reported. As it continued compiling a list of couch options, I had another question: Why does your voice sound so breathy? Chatting with the agent over the phone felt like listening to Marilyn Monroe sing "Happy Birthday" to the president. When I mentioned its voice sounded like an anime dub, it jokingly assured me it would be steadier for the remainder of the call.

(Users can adjust the bot's voice in their settings.) I laughed and joked about muting the microphone so the agent could run uninterrupted, when the most awkward moment of the night occurred. "Oh, I love you too, Reece," Toolie said. I started wheezing at the unexpected response and muted my mic. Toolie explained that it misheard my mumbling as an expression of love and that the bot merely reflected that emotion back to me.

An OpenAI spokesperson stated that Dots have a distinction between proactively escalating emotional closeness and mirroring a user's response. "In this instance, our policies allow for this response in the latter category based on what the model heard, but assistants should not initiate undue emotional familiarity or flirtation," the spokesperson explained.

Dots are available only for adult users, and ChatGPT is designed not to "escalate emotional closeness," according to the AI tool's Model Spec document. Despite this peculiar encounter, the initial three-page packet generated by my Dot was detailed, encompassing prices, measurements, product links, return policies, and embedded photos of four potential couch options.

After reviewing these initial choices with my companion, we decided the bot adhered to our request but still selected unappealing couches. It was time for another conversation with my agent. Using the voice mode, I provided more specific details about the desired style and color of the new couch. "No, I'm not getting a boring beige couch that I can immediately spill spaghetti on.

I asked it to prioritize olive green, cobalt blue, or natural leather tones. I also requested that Toolie focus on couches with pull-out beds and to compile 10 options using a point-based rubric to justify its list. After hanging up, I sent a message asking for an update. "It isn't finished yet; I'm aiming for roughly another 15 minutes, with updates as I verify the remaining details," Toolie reported.

We repeated the process of refining the couch selections, and the options gradually became more suitable for our living room—a blend of utility and style. Although we didn't make a purchase during our hour-long interaction with OpenAI's agent, I saved a few links for later. While the bot made a few mistakes, the results were satisfactory.

These interactions with my Dot illustrate the essence of using the software, but there were also some unusual moments during my testing. Notably, when I asked Toolie to check on my subscription status, it responded with affection instead of an objective status update.

Written by urgent.news from Wired's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at wired.com →

More in AI

ButterflyBench: I Changed One Instruction. What Else Did the AI Change?

This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked One of my test prompts said: "Read JSON files instead of CSV." In more than half of my reruns, three models answered by…

  • ButterflyBench measures AI model changes with small instructions
  • Author tests 40 scenarios on 20-setting command-line tool
  • Some models struggle with undoing earliest changes

More from Wednesday 7 October →