Urgent.News

What's breaking now, across thousands of outlets.

AI

Six Months of Writing Code Exclusively with Agents

In February, a writer made a rule for themselves: no more writing code by hand. They had been doing so for six months prior, relying on their knowledge of the system to quickly build features and fix bugs. However, as more people contributed to the codebase, the writer spent more time reading changes and less time actually writing code. Typing speed and the complexity of changes were the main challenges in this approach.

The writer tried using AI agents to help with the coding process. Tools like Copilot autocomplete and Cursor's tab complete helped initially, but as the models improved, they could generate complete changes to multiple files at once. This reduced the amount of typing needed, but the writer still had to review and edit the generated code to ensure it matched their desired state. Agents would often make incorrect changes, so the writer needed to manually edit some of the generated code.

In early 2024, the AI models, such as GPT-5.3 and Opus 4.6, became significantly better at handling larger changes with less guidance. This allowed the writer to stop writing code by hand completely. If an agent got stuck, the writer was not allowed to finish the code themselves; they had to figure out what the agent was missing and fix it instead.

The writer emphasized that they didn't learn coding from reading about it, but through writing a lot of code, running it, seeing it fail, fixing it, and doing it again. They viewed AI agents as software that needed to be used for real work, where they could see where they failed, change the prompts, tools, or environment, and try again.

Despite the benefits of using AI agents, the writer encountered challenges with parallelism. Running multiple agents at once led to conflicts in the same files and Git state, as well as issues with dependencies, ports, and processes. To mitigate these problems, the writer asked colleagues for workarounds, such as using worktrees for each agent or creating ephemeral databases and ports.

However, these solutions had limitations, as the agents still shared databases, ports, processes, and other machine resources. The writer also had to coordinate when each agent could test, push, or deploy, adding to the complexity.

To address these challenges, the writer moved their agents off their laptop and onto separate machines, using exe.dev's Linux VMs that could be brought up in a few seconds with SSH and HTTPS already set up. Each task was assigned its own machine, allowing the writer to walk away from their laptop while the work continued. The writer created a startup script using Claude to automate the setup of a complete development environment on a fresh VM.

This script installed toolchains, cloned repositories, configured Claude Code and Codex, and ensured the environment was ready for the agent to work.

While the agent boxes worked well, the writer still faced the issue of each agent having its own tmux session. To streamline the process further, the writer ended up with separate tmux sessions for each agent.

Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at blog.exe.dev →

More in AI

Your local LLM app needs guardrails before it needs prompts

Most local-LLM tutorials start with the fun part: the prompt. After running a fleet of autonomous agents on local models 24/7 and logging every failure — the ledger now holds over eight thousand…

  • Scaffold includes output contract with minChars, must, and mustNot clauses.
  • Retry mechanism feeds failures back into next prompt after rejection.
  • Approval queue requires human approval before output leaves the app.

Beyond the Prompt: Why the Hermes Agent Is the Self-Improving AI We Actually Needed

We have all hit the "chatbot wall." You open a clean UI, paste a massive block of context, get a decent response, and then close the tab. The next day? You start completely from scratch.

  • Hermes Agent is self-improving AI developed by Nous Research
  • Unlike traditional chatbots, Hermes learns, plans, and adapts to user's work style
  • Hermes' persistent memory system retains context across sessions and platforms

I Tested GLM-5.3-Flash and Qwen3.8-Flash on 24 Real Tasks

I test-ran both of this week's open-weight flash models against 24 small, real workloads from an actual product stack — structured extraction, SEO metadata, and code fixes — and graded everything…

  • GLM-5.3-Flash and Qwen3.8-Flash tested on 24 real tasks
  • Both models achieved 90% success in structured extraction tasks
  • GLM-5.3-Flash better for code generation, Qwen3.8-Flash for SEO metadata

More from Thursday 27 August →