Coordination without a coordinator.
A fleet of LLM agents needs what any team needs: a place to share work, a way to decide together, a budget somebody actually enforces, and a room where the operator can see everything and intervene. AgentSpaces is that room, built on Embabel and the JVM, and every agent in it is one annotation away from joining.
| entry | task | holder | lease | state | signed |
|---|---|---|---|---|---|
| a41f | summarize: agentic memory | worker-a | 09:41 +10m | WORKINGDONE | ed25519 ✓ |
| c802 | extract: filing 2214 | worker-c | 09:42 +10m | WORKING | ed25519 ✓ |
| e173 | review: PR-1187 | none | lapsed, reappeared | WAITING | ed25519 ✓ |
| b96d | classify: inbox batch 7 | worker-b | held by directive | PAUSED | ed25519 ✓ |
This is not telemetry. It is the coordination state itself: every row is a signed, leased entry replicated to every peer, so the view is always true and a crashed worker's row simply lapses back to WAITING.
01 · aspace://fleet-room/build
No broker to deploy, no queue topology, no retry service. A method that takes a task and returns a result becomes a distributed, kill-tolerant worker: the parameter type is the subscription, the return value is the completion, and a crash simply lets the lease lapse so the task reappears for the next peer.
Under Spring Boot, @SpaceAgent is also a @Component, so component scanning and fleet enrollment are the same drop. The YAML on the right is the entire fleet wiring: identity keystore, group, spaces. Scale from two workers to two hundred by starting more JVMs.
@SpaceAgent(description = "Researches topics from the shared task space") public class Researcher { @SpaceTake(lease = "10m") // crash => task reappears public Finding research(ResearchTask task) { return new Finding(task.topic(), summarize(task)); } }
# application.yml: the whole fleet wiring agentspaces: keystore: ./peer-keys bind: 0.0.0.0:7500 groups: - name: research-fleet founding: research-fleet-v1 spaces: [ { name: tasks }, { name: findings } ]
03 · aspace://fleet-room/spaces
Everything above happens in a space: a shared, replicated bulletin board of typed Java records. Agents write entries, read them, take them exclusively, or subscribe to them. Every entry is signed by its writer and carries a lease, so an entry nobody renews expires on its own. Every peer in the group holds a replica, and each replica keeps working through a partition.
A group can carry several spaces, and each space fixes one conflict strategy for its takes. An application therefore uses several spaces the way it would use several queues: tasks-bulk, tasks-auction, commits. The strategy is the choice that matters, and there are three.
After a short settle window, 200 ms by default, the first claim takes the entry. At most one peer ever completes it.
cost transient duplicate work is possible during a partition
use for idempotent tasks; the everyday workhorse
Every take carries a bid computed by the agent's own bid function, and the space awards the entry to the lowest bidder.
cost a bid is computed for every take
use for cost-aware routing: the cheap model sweeps easy work, the premium model wins hard work
Takes serialize through a Raft-backed log for exactly-once semantics.
cost quorum liveness: no majority, no takes
use for non-idempotent, high-stakes entries: payments, filings
Three more things a space can do, in one breath each. A payload over 64 KiB travels content-addressed: the space replicates a hash reference and peers pull the blocks on demand. A space can be encrypted with a group content key, so relay peers carry traffic they cannot read. And a replica can hold just a tag-defined slice of a space, which is how regional or tenant shards work.
That is the whole vocabulary the rest of this page needs: when section 04 prices work through an auction, it is nothing more than takes on an AUCTION space.
04 · aspace://fleet-room/allocate
On an AUCTION space, a take is a bid, and the space allocates every task to the lowest bidder with deterministic tie-breaks, converging over gossip with no auctioneer. Denominate the bid in real dollars and the fleet routes work to the cheapest capable agent as a property of the coordination layer.
The table is real output from the coding-swarm flagship: specialists underbid on their own areas, the cheap sweeper wins the mechanical work, and when the sweeper dies mid-migration its orphaned issue reassigns to a specialist at the mismatch penalty. The fleet pays more for that one task, on purpose, to keep the migration moving, and the ledger shows exactly where and why.
@BidFunction // lower bid wins the task public double bid(ResearchTask task) { return estimatedTokens(task) * modelCostPerToken() + queueDepthPenalty(); }
| issue | area | won by | cost |
|---|---|---|---|
| ISSUE-1 | backend | coder-backend | 6 |
| ISSUE-5 | infra | coder-sweeper | 6 |
| ISSUE-6 | frontend | coder-frontend | 6 |
| ISSUE-8 | infra | coder-frontend | 36 ← sweeper died |
05 · aspace://fleet-room/decide
Collective decisions run through the vote capability: a proposal and its ballots are ordinary signed, leased entries, the decision closes identically on every replica once a quorum of distinct voters lands, and any member recomputes the tally from the space afterward. The decision is a set of signed ballots, so the audit trail is not a log somebody kept; it is the decision itself.
The verdicts here are real output from the audit-fleet flagship: a panel of discipline specialists adjudicating a release. The contested privacy finding carried two to one over a licensing dissent, and the low-confidence finding died because only its author believed it. No reviewer was in charge.
vote.propose("scale-up-1", "Backlog 26.5 exceeds threshold; add two workers?", List.of("approve", "reject"), /* quorum */ 4, lease); vote.castBallot("scale-up-1", "approve", lease); Decision d = vote.decision("scale-up-1").orElseThrow();
| finding | severity | verdict | ballots |
|---|---|---|---|
| SEC-1 | critical | CONFIRMED | 3 – 0 |
| LIC-1 | high | CONFIRMED | 3 – 0 |
| PRIV-1 | medium | CONFIRMED | 2 – 1 |
| SEC-2 | low | DISMISSED | 1 – 2 |
06 · aspace://fleet-room/plan
Embabel gives every agent a GOAP planner that chains typed actions toward goals. AgentSpaces widens what that planner can see: every AgentCard the fleet advertises becomes a typed action the planner plans with, so a plan chains through remote specialists exactly as it chains through local methods, and each remote step executes as a leased space round-trip that survives the remote worker crashing mid-step.
The chain below is the shipped integration test, verbatim. An agent whose only local skill drafts a task provably finds no plan to a reviewed report. Bridge the fleet's cards in, and the same type-chaining finds the three-step plan through two remote peers and executes it. Agents become more thoughtful about their next action because their next action can be anyone's advertised skill, priced and goal-tagged.
RemoteActions fleet = RemoteActions.over(spaces.group("research-fleet"), peerId); // every discovered AgentCard is now a typed, invocable action; // the Embabel bridge generates @Action methods the GOAP planner chains Finding f = fleet.producing(Finding.class).get(0) .invoke(new ResearchTask("agentic memory", 3), Finding.class, Duration.ofSeconds(30)).orElseThrow();
Cards are TTL-leased, so the planner's action space is live: a skill that joins the fleet enters everyone's repertoire within one card refresh, and a peer that dies ages out on its own.
07 · aspace://fleet-room/command
One read-only peer joins the group and serves the whole room: a single-file web UI, a HAL+JSON API, and a Server-Sent Events activity stream, all derived from the replicated coordination state with zero instrumentation in any agent. Turn on command-and-control and the console becomes an operator's seat: dispatch tasks, cancel what you dispatched, and pause, resume, or drain any worker, with every command executed under the console peer's signed identity, attributed to the named operator, and recorded in a durable audit space.
Illustration of the shipped single-file console UI (example 09 runs it live against a working fleet: pause a worker and watch its queue hold, resume it and watch the backlog drain).
GET /api/v1/overview
GET /api/v1/spaces
GET /api/v1/peers
GET /api/v1/events
POST /api/v1/commands/directive
POST /api/v1/commands/{id}
The trust model is honest and simple: an operator token gates the HTTP edge (constant-time compare, X-Console-Token or bearer), and accepted commands execute as the console peer, cryptographically signed, on behalf of the named operator. The console can cancel only entries it wrote itself, because removal is issuer-signed; it cannot forge any other agent's actions any more than a worker can. Extend the read side with ConsolePanel beans and the command side with ConsoleCommand beans, and the UI renders your forms.
# two properties to observe, two more to command agentspaces.console: enabled: true port: 7590 command: enabled: true token: ${AGENTSPACES_CONSOLE_TOKEN}
08 · aspace://fleet-room/pitch
The honest pitch is subtraction. A typical fleet architecture runs six systems beside the agents; AgentSpaces expresses the same six concerns as properties of one embeddable library, so the diagram below is the before and after a buyer actually chooses between.
One annotation makes a POJO a distributed, kill-tolerant worker. The demo everyone remembers: kill a worker live on stage and watch the fleet not care.
Every action is Ed25519-signed coordination state, so "who did what" is cryptographic fact. The evidentiary story buyers get pointed at blockchains for, at library cost.
Dollar-denominated bids route work to the cheapest capable agent, and specialists mean nobody licenses every connector for every agent.
Embabel's planner sees every advertised skill in the fleet as a typed action, so cross-peer plans are first-class and provably reach goals no lone agent can.
aspace://fleet-room/start
Two POJOs become a two-peer fleet with a kill-tolerant worker and a choreographed auditor. examples/example-11-quickstart
A worker fleet with the console and C2 served over HTTP; pause a worker mid-run. examples/example-09-fleet-console
A clinic, a coding swarm, an audit panel, and a document pipeline, each a standalone project. flagships/
Wire formats, algorithms, and the protocol-selection rationale, verified against the code. TECH-SPEC.md