Urgent.News

What's breaking now, across thousands of outlets.

AI

I Built a Village of Six Agents You Can Actually Inspect

An agent claims a task. Another agent needs the same resources. A third is resting. Who gets the job, and can you explain the decision afterward? That is the kind of question I am exploring with Settlement , a village strategy game and an inspectable simulation lab. With Hacktoberfest focusing on open-source AI this October , it feels like a useful time to share a project where the behavior is…

In a village strategy game called Settlement, a team of six agents cooperates to achieve a common objective, such as preparing for a raid or restocking supplies. The game operates as an inspectable simulation lab, allowing users to observe the agents' behavior and decision-making process.

When an agent is assigned a task, another agent may require the same resources. A third agent may simply be resting. The game's planner divides the objective into individual jobs, with each job assigned to a single agent. The agents then walk to their designated workplaces, gather resources, and coordinate recruitment as needed.

If a resident chooses to rest, their work can be released for another resident to take over. When resources or workplaces change, the agents reassess their tasks. The game's inspector feature provides visibility into current jobs, blocked work, spending limits, and recent memories of the agents.

The game uses deterministic game AI, meaning it does not understand arbitrary natural-language goals, train neural networks, or call a model behind the scenes. Optional Claude proposals are part of explicitly configured local-worker research experiments. Keeping these paths separate allows for clearer measurement of results.

Settlement enforces constraints on actions, ensuring that they validate funds, capacity, and the objective's spending allowance before altering the treasury or army. To try the game locally, clone the repository from GitHub, navigate to the project directory, install dependencies, and run the development server. Open the game at http://localhost:5173, choose Village orders → Prepare for a raid, and then View plan → Shared jobs to inspect one resident before and after giving it a rest.

To gain a deeper understanding of the game's behavior, ask three questions: Did the job change owners when the resident rested? Can you identify why the replacement took the task? Did the action respect the objective's budget? The research view offers matched seeds, JSON/CSV exports, and replay verification. The game's authored scenarios serve as a bounded testbed, providing results that should not be considered a general AI safety benchmark.

Contributions to the project could explore richer objectives, recovery behavior, or more adversarial and honest-control scenarios.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Opus 5.5 loves to tell you ‘this matters’ (and other AI writing tells)

Opus 5.5’s biggest tell is the word “dependable,” which pops up 23 times more often than in human samples.

  • "Dependable" appears 23 times more often in Claude Opus 5.5's writing than in humans
  • "This matters" occurs 116 times more frequently in Opus 5.5 than in human writing
  • Astra model emphasizes "another dimension" and uses "may provide" or "can provide" hedging

An Analytics Agent's Permissions Should Survive a Bad Prompt

An analytics agent receives a hostile instruction: return every customer's unmasked payment identifier. The key question is what the system permits that agent to read—even if the model decides to…

  • Agent operates under service principal with Unity Catalog permissions
  • Permissions depend on correctly configured identities, grants, masks, filters
  • Testing protected tables under actual identity ensures thorough access control

I gave my coding agent a sense of taste. It picks restaurants from my music.

My coding agent can refactor a thousand lines without breaking a sweat, but ask it "where should I take a date Friday night" and it's useless. It knows my code. It knows nothing about my taste.

  • Qloo coding agent gains taste sense
  • Qloo Taste API connects 250M cultural entities
  • Agent recommends restaurants based on music preferences

More from Thursday 1 October →