aspace://guide/quick-start

Quick start

Two plain classes become a two-peer fleet: a kill-tolerant worker and an auditor that reacts to every result. Nothing here deploys a broker, a registry, or a scheduler; the coordination fabric is a library inside your application.

1. Add the starter

<!-- pom.xml -->
<dependency>
  <groupId>ai.badmonkey.agentspaces</groupId>
  <artifactId>agentspaces-spring-boot-starter</artifactId>
  <version>0.1.0-SNAPSHOT</version>
</dependency>

2. Describe the fleet, once

The starter reads one YAML block: an identity keystore, a group every member derives the same self-certifying id from, and the spaces the group shares. Start a second JVM with the same block and a seed address, and you have a fleet.

# application.yml
agentspaces:
  keystore: ./peer-keys
  bind: 0.0.0.0:7500
  groups:
    - name: orders-fleet
      founding: orders-fleet-v1
      seeds: ["hq.example:7500"]     # omit on the first peer
      spaces: [ { name: work } ]

3. Write the agents

The worker takes orders under a lease and returns shipments; the auditor reacts to every shipment and returns a receipt. Both are complete as shown.

quickstart — orders-fleet
@SpaceAgent(description = "Ships orders from the shared space")
public class Fulfiller {

    @SpaceRef
    private Space space;                    // injected at bind time

    @SpaceTake(lease = "10m")              // crash => order reappears
    public Shipment ship(Order order) {
        space.write(new Progress(order.orderId(), "picking"),
                    Lease.of(Duration.ofMinutes(5)));
        return new Shipment(order.orderId(), order.item(), "fulfiller");
    }
}
✓ binds on startup · worker loop + subscription running · AgentCard published

4. Put work in

spaces.group("orders-fleet").space("work")
      .write(new Order("ord-1", "kite"), Lease.of(Duration.ofMinutes(10)));

What you now have: the order replicates to every peer, one fulfiller wins the take fleet-wide, the shipment and receipt appear as signed entries, and if you kill the working JVM mid-task the lease lapses and the order reappears for another. Run examples/example-11-quickstart to watch exactly this over two TCP peers.

No Spring requiredThe same annotations bind anywhere: create an AgentBinder, register spaces with binder.space("work", space), and call binder.bind(new Fulfiller()). The starter only automates what plain Java does in three lines.
aspace://guide/annotations · chapter 1

Annotations

The annotation model is the top of the stack and the whole API most applications need. Six annotations cover identity, work, choreography, economics, access, and capability services; each maps to protocol behavior the layers below carry for you.

They go on ordinary objects. There is no base class to extend, no interface to implement, and no container of ours to run in: the binder reflects over an instance you hand it, wires whatever it finds, and leaves the rest of the class alone. Constructors, other Spring beans, JPA repositories, your own fields — all untouched. That matters for how you organize a fleet, so each annotation below says plainly where it belongs and what else is allowed to sit beside it.

Where these liveOn any class you can get an instance of. Under Spring, annotate the bean (@SpaceAgent is @Component + @AgentSpec) and the post-processor enrolls it after initialization. Without Spring, call binder.bind(instance) or spaces.bind(instance) yourself. Either way you construct the object normally first — the binder never instantiates anything.

@SpaceAgent and @AgentSpec

@AgentSpec declares an agent's identity on any class, framework or not. Under Spring, prefer the @SpaceAgent stereotype: it is both a @Component and an @AgentSpec, so component scanning and fleet enrollment are one drop. Binding publishes an AgentCard built from the identity plus the bound method signatures, so other peers can discover what this agent consumes and produces with no extra code.

@SpaceAgent(name = "researcher", description = "Researches topics",
            goals = {"answer research questions"})
public class Researcher { ... }
attributedefaultmeaning
namedecapitalized class namethe agent's local name; the AgentId is the peer id plus this name
description""for humans and semantic discovery; also feeds Embabel planning (chapter 3)
goals{}goal statements published on the card
Composed stereotypesThe binder resolves @AgentSpec through one level of meta-annotation. Annotate your own annotation with @AgentSpec, give it name/description/goals attributes, and it works exactly like @SpaceAgent does.

@SpaceTake: the kill-tolerant worker

Turns a method into a leased take loop. The parameter type is the template; a normal return completes the take and writes the result atomically; a thrown exception skips completion so the lease lapses and the task reappears for another worker. The failure path is the happy path with nothing added.

// agentspaces-partybus: one of eight scouts, each a class with one take method
@AgentSpec(name = "flight-scout",
        description = "Finds flights and rail for a trip's outbound and return",
        goals = {"find transport for a trip"})
public static final class FlightScout {
    private final Sources sources;           // an ordinary constructor dep

    public FlightScout(Sources sources) {
        this.sources = sources;
    }

    @SpaceTake(space = "research", lease = "2m", pollTimeout = "500ms")
    public FlightOptions scout(FlightRequest request) {
        List<Link> web = sources.search("flights " + request.from()
                + " to " + request.to(), 8);
        return sources.structured("flights", prompt(request, web),
                FlightOptions.class);
    }
}

Reach for @SpaceTake whenever work is a queue of units that exactly one worker should do: research requests, document conversions, booking orders, anything you would otherwise put on a job queue. Scale is horizontal and needs no configuration — start a second process with the same class and the same queue drains twice as fast, because both are racing for the same entries. The lease is the whole failure story: kill a scout mid-search and its TAKE lease lapses, the request reappears, and a sibling picks it up.

Where it livesOne method on a class that may hold anything else. The parameter type is the queue, so one class can carry several @SpaceTake methods over different entry types, or you can give each worker kind its own class, as the partybus scouts do. Constructor dependencies are yours to supply: the binder binds an instance you already built.
attributedefaultmeaning
space"" (inferred)the space to take from; empty means the group's sole registered space
group"" (any)the group this binding belongs to, by name or GroupId; empty means any group whose spaces satisfy it
lease"10m"the TAKE lease; how long a crash takes to surface as reappearance
pollTimeout"1s"how long each poll waits before looping
resultLease"1h"write lease for the returned result entry
resultSpace"" (same space)where the result entry is written
Keep in mindThe method body runs on the binder's worker thread for that method, one entry at a time — so a slow method is backpressure, not a queue that grows behind you; add another process to go faster. Return null (or declare void) to complete the take and write nothing. Throw, and the binder deliberately does not complete: the lease lapses and the entry comes back, which means a task that always throws will be retried forever. Catch what you can handle and write a failure entry for what you cannot. And keep the method's own work shorter than lease, or the entry reappears while you are still working on it and two workers finish the same task.

@SpaceNotify: choreography as a returning method

Invoked for every matching entry written to the space, without consuming it, and a non-void return value is written back as the next entry in the flow. React to X, produce Y; Y is some other agent's cue. The binder carries the discipline so the method body never has to: delivery is deduplicated per entry (at-least-once becomes effectively once per bound agent), the reaction runs on its own virtual thread so slow work never holds the fabric's delivery thread, and a null return writes nothing.

// agentspaces-partybus: the dispatcher reacts to one brief by fanning out
// a dozen requests. It writes requests and never answers them.
@AgentSpec(name = "dispatcher",
        description = "Turns a travel brief into the fleet's research requests",
        goals = {"research a trip"})
public final class Dispatcher {

    @SpaceRef("research") Space research;      // the bag the scouts take from
    @SpaceRef("trip")     Space trip;          // where briefs and events arrive
    @SpaceRef("slots")    Space slots;         // the day auction

    @SpaceNotify(space = "trip", lease = "12h")
    public void dispatch(TravelBrief brief) {
        String id = brief.tripId();
        research.write(new FlightRequest(id, brief.from(), brief.to()), REQUEST_LEASE);
        research.write(new HotelRequest(id, brief.to(), nightly), REQUEST_LEASE);
        research.write(new CarRequest(id, brief.to()), REQUEST_LEASE);
        for (String date : brief.days()) {
            slots.write(new DaySlot(id, date, brief.to()), REQUEST_LEASE);
        }
    }
}

Reach for @SpaceNotify when the trigger is "something appeared" rather than "there is work to claim". Nothing is consumed, so every interested agent sees the same entry — which is exactly what you want for fan-out, for derived records, and for side effects like audit and notification. Returning a value is the compact form of a pipeline: an agent that returns Y from an X has declared one edge of a flow, and the next agent's @SpaceNotify on Y is the next edge. Nobody wires the graph; it exists because of what each agent consumes and produces.

The example above shows both halves of the common shape: react with @SpaceNotify, then write through @SpaceRef handles when one reaction needs to produce many entries, or entries in other spaces, rather than the single entry a return value can express.

Where it livesSame freedom as @SpaceTake, and the two mix in one class: the partybus planners each carry a @BidFunction, a @SpaceTake, and a @SpaceRef together. A class with only @SpaceNotify methods is a perfectly good agent — the pet-clinic flagship coordinates entirely this way, with no take anywhere.
attributedefaultmeaning
space"" (inferred)the space to watch
group"" (any)the group this binding belongs to, by name or GroupId
lease"1h"the subscription lease; the binder renews it at half-lease intervals while bound, so it bounds how long a stale subscription outlives a crashed agent
resultSpace"" (same space)where a returned entry is written
resultLease"1h"write lease for a returned entry
Keep in mindEvery bound agent sees every matching entry, including entries this same process wrote — so a method that reacts to X by writing another X is an infinite loop. Make the reaction produce a different type, or filter on a field. Reactions run on their own virtual thread per entry, so two entries can be in flight at once and your method must be safe to re-enter; the binder deduplicates redeliveries of the same entry, not concurrent different ones. A reaction that throws is logged and dropped: there is no lease to lapse and nothing will retry it, which is the opposite of @SpaceTake and the single most common surprise when moving logic between the two.

@BidFunction: the price of work

On an AUCTION space, every take carries a bid and the lowest bid wins the entry. The bid is a plain method over the entry, so cost models are ordinary code: denominate in dollars, tokens, or queue depth, and let FinOps tune it.

// agentspaces-partybus: one planner. Bid, worker, and space handle all in
// one class -- the bid decides whether this planner gets the day at all,
// and the take method is what it does once it has won.
@AgentSpec(name = "dining-planner",
        description = "Plans a day around the table: markets, long lunches",
        goals = {"fill a day with food"})
public static final class DiningPlanner {

    @SpaceRef("research")
    Space research;

    @BidFunction(space = "slots")               // lower bid wins the day
    public double bid(DaySlot slot) {
        double bid = 90;
        if (interested(slot, "food", "wine", "eat")) {
            bid -= 35;                          // this is our kind of day
        }
        if (slot.pace().equals("REST")) {
            bid += 20;                          // ... but not on a rest day
        }
        return bid + slot.round() * 5;          // a strike barely touches a lunch
    }

    @SpaceTake(space = "slots", lease = "1m", pollTimeout = "500ms")
    public ScheduledActivity plan(DaySlot slot) {
        String table = research.readAll(Template.of(RestaurantOptions.class)
                        .where("tripId", eq(slot.tripId())), 20).stream()
                .flatMap(r -> r.restaurants().stream()).map(Restaurant::name)
                .findFirst().orElse("a table the scouts are still finding");
        return new ScheduledActivity(slot.tripId(), slot.date(), "DINING",
                "Market morning, long lunch, dinner at " + table, bid(slot));
    }
}

This is the shape worth studying, because it answers the organizing question directly: the bid function, the worker, and the space handle live in the same object. That is deliberate. A bid is a statement about what this particular agent would charge, so it needs the same state the work needs — the model it would call, the queue it is already sitting on, the region it runs in. Splitting them apart would mean keeping two things in sync for no gain.

Reach for AUCTION when workers are not interchangeable and you want the fleet to sort that out without a router. Several planners bid on every day of a trip; the one whose day it should be bids lowest and wins, and nothing anywhere contains a rule saying who plans what. Change a bid function and you have re-tuned the whole allocation. The same shape does cost-aware model routing (cheap capacity sweeps easy work, the expensive model wins the hard entries), locality routing, and load shedding by bidding your own queue depth.

Where it livesIn the same class as the @SpaceTake it prices, on an AUCTION space. One bid function per space per bean; the binder wires bids before starting worker loops, so a bid is always in place before the first claim. A bean with a bid function but no take method is legal and useless — it prices work it will never do.
Keep in mindLower wins, and the number is yours to denominate — dollars, tokens, queue depth, estimated seconds — as long as every bidder on that space uses the same unit; they are compared directly. The function must be pure and quick: it runs inside the claim path, on candidate entries, possibly many times. Do not call a model, hit a database, or block in it. Only spaces built with ConflictStrategyType.AUCTION consult bids at all; on a LEASE_RACE space a @BidFunction is silently never called, which is a common reason a "broken auction" turns out to be a misconfigured space. And a space consults one cost function: two agents with a @BidFunction on the same space on one peer are refused at bind time, by name, rather than the second silently replacing the first — example-04-auction shows the refusal and then the shape that works, one bidder per peer.

@SpaceRef: a handle when you need one

Injects a Space into a field at bind time, for mid-method writes such as progress markers and side outputs. Under Spring, @Scheduled plus a @SpaceRef field is the idiomatic periodic agent; AgentSpaces ships no scheduler of its own. What arrives is the registered space itself — or, when the fleet runs with agentspaces.identity.agent-keys=subordinate, this agent's view of it, so everything the agent writes, takes, or completes through the field is signed with the agent's own certified key and attributed to the agent, not merely to its host peer (chapter 6). The field's type and your code do not change either way.

@SpaceRef("findings")
private Space findings;                  // injected at bind time

// Several are normal: the partybus dispatcher holds three, one per space
@SpaceRef("research") Space research;
@SpaceRef("trip")     Space trip;
@SpaceRef("slots")    Space slots;

// And the idiomatic periodic agent under Spring -- our scheduler is Spring's
@Component
@SpaceAgent(description = "Publishes a backlog reading every thirty seconds")
public class BacklogReporter {

    @SpaceRef("metrics")
    private Space metrics;

    @Scheduled(fixedDelay = 30_000)
    public void report() {
        metrics.write(new Backlog(host(), depth()), Lease.of(TWO_MINUTES));
    }
}

Use it whenever one invocation has to produce more than the single entry a return value can express: progress markers mid-task, fan-out like the dispatcher's, side outputs into a different space, or reads that inform the work (the dining planner reads what the scouts found before it writes a schedule). The field is a fully-featured Space — every verb of chapter 4, not a write-only sink.

It is also the answer to "how do I get a space into code that is not a bound method". A @Scheduled method, an HTTP controller method, a message-listener callback: annotate the field, and as long as the enclosing bean is enrolled, the handle is there. AgentSpaces ships no scheduler, no HTTP layer, and no inbox of its own precisely because your framework already has them.

Where it livesA field on the agent class, or on any superclass of it — the binder walks the hierarchy. The field type must be exactly Space; anything else fails at bind time rather than at first use. Fields need not be public. Injection happens at bind time, so the handle is null in the constructor and in any @PostConstruct: do first-write-on-startup work from a @SpaceNotify, a scheduled method, or an ApplicationReadyEvent listener, never in the constructor.

@Ballot: one signed vote per proposal

The first of the Layer 4 annotations, and the shape the others follow: the method's parameter is the cue and its return value is what the binder does with it. Here the cue is the vote's own Proposal entry and the return is the option this agent backs; the binder casts the ballot through the group's vote capability, as this agent, exactly once per proposal, and a null return abstains. Because the cue is the proposal record, the ballot can never be refused as "unknown proposal" — the two-entry race every hand-written panelist had to defend against does not exist.

// flagships/agentspaces-audit-fleet -- AuditFleet.Panelist
@Ballot(space = VOTES, lease = "2h")
public String judge(VoteCapability.Proposal proposal) {
    Finding finding = peer.audit().read(Template.of(Finding.class)
            .where("findingId", eq(proposal.proposalId())), Duration.ofSeconds(10)).orElse(null);
    if (finding == null) {
        return null;                                   // nothing on the record: abstain
    }
    double weighted = finding.confidence() * trust(peer.discipline(), finding.discipline());
    return weighted >= CONFIRM_THRESHOLD ? "CONFIRMED" : "DISMISSED";
}

Attributes: space (the vote space; empty resolves to the sole registered vote), prefix (only proposals whose id starts with it — one vote space usually carries several decisions, the Party Bus keys them advisory: and next:), lease for the ballot, and group. Under agent-keys=subordinate the ballot is the agent's own attested record, so several agents on one peer are several voters wherever the authorizer counts per agent (chapter 6).

Where it livesAny bound class. The space needs a vote capability registered for it — group.provide(vote) does it (and advertises the capability), group.vote("votes", vote) registers without advertising, binder.vote(...) is the hand-built form — and binding fails fast otherwise, naming that call. The option must be one of the proposal's; casting one that is not throws exactly as castBallot does.

@OnDecision: react once when a vote closes

The pair of @Ballot. The method takes the Decision and runs exactly once per proposal, the first time this replica's tally meets the quorum; a non-null return is written as the next entry in the flow, like @SpaceNotify. Underneath, the binder watches ballots and proposals in the vote space and recomputes the decision on each, so a replica that learns of a vote late still fires once, on its first qualifying ballot. This replaces the decision() poll loop that every consumer of the vote used to write.

// examples/example-05-quorum -- QuorumFleet.Supervisor
@OnDecision(space = "votes", prefix = "scale", resultLease = "10m")
public ScalingDecision record(VoteCapability.Decision decision) {
    return new ScalingDecision(decision.proposalId(), decision.winner(),
            decision.tally().toString(), decision.granularity().name());   // which counting rule closed it
}
Where it livesThe same class as the @Ballot, or a different one: the audit flagship's lead reacts to closed votes and never casts. resultSpace defaults to the vote space; the Party Bus's safety lead writes the advisory into trip. Plain Java without the binder has the same hook, vote.onDecision(proposals, listener, lease), which returns a leased Subscription.

@OrderedTake: the exactly-once worker

@SpaceTake's contract — take, act, complete-or-lapse — with the take routed through the space's ordered-log coordinator instead of the space's own claim race, so the log's commit order decides every take identically on every member and each entry is completed at most once, fleet-wide. The binder owns the loop the coordinator asks of a caller: a take lost to a leader election is simply resubmitted on the next poll. This closes what the space chapter used to call "the one place the annotations cannot reach".

// examples/example-12-exactly-once-desk -- ExactlyOnceDesk.Clerk
@OrderedTake(space = PAYMENTS, lease = "30s", pollTimeout = "2s", resultLease = "1h")
public PaymentReceipt confirm(PaymentOrder order) {
    return new PaymentReceipt(order.orderId(), order.payee(), order.cents(), peer.name(),
            peer.raft().commitIndex());                        // completes the take with the receipt, once
}

// ExactlyOnceDesk.startClerk(...): the coordinator and its log, built together and registered
OrderedTakes ordered = OrderedTakes.over(new CapabilityPipes(runtime, codec), codec,
        identity, name, members, InstantSource.system(), seed, payments);
group.ordered(PAYMENTS, ordered);      // the log rides the peer tick; role=leader is discoverable
Where it livesAny bound class, on a peer whose group has a coordinator registered for the space (group.ordered(space, coordinator)); binding fails fast otherwise, and also when the coordinator takes as a different agent than the space handle writes as. The member set is still fixed and still yours to supply — that has not changed — but the AtomicReference cycle that wired the log to its coordinator is gone: OrderedTakes.over(...) does it.

@CapabilityRef: a typed client, injected

The sibling of @SpaceRef for Layer 4: a field of a typed client type — VoteClient, AggregateClient, SemanticClient, or anything registered with binder.client(type, instance) — is injected at bind time from the group's typed clients. Constructor injection still works and is still fine; this is for when the agent should not have to be handed its capabilities by whoever constructs it.

@CapabilityRef VoteClient votes;
@CapabilityRef AggregateClient aggregate;
@SpaceRef("trip") Space trip;

// A push-sum contribution is a return value, not a call: the binder starts the epoch.
@SpaceNotify(space = "trip", lease = "12h")
public Contribution feel(TripDay day) {
    return Contribution.to(Itinerary.energyEpoch(day.tripId(), day.date()), energyOn(day));
}
Where it livesA field on the agent class or a superclass. Binding fails fast when no client of the field's type can be resolved, naming the registration that would supply it. Returning a Contribution from a @SpaceNotify, @SpaceTake, or @OnDecision method needs the peer's aggregate registered (group.provide(aggregate)), and the fleet's estimate is read back with AggregateClient.awaitEstimate(epoch, timeout) — there is no manual tick anywhere.

@ProvidesCapability: serving a capability

Marks a CapabilityProvider implementation as a service the fleet should advertise. Bound through the facade, it is registered on the group's CapabilityRuntime, which starts it, signs and publishes its CapabilityAdvertisement, and keeps that advertisement fresh on every card refresh — the same TTL-leased freshness every other card gets. Consumers reach it through a typed client (chapter 5).

// A capability bean: wrap a shipped capability (or your own) and advertise it.
// Under Spring the facade bean supplies the space the capability runs over.
@Component
@ProvidesCapability(VoteCapability.TYPE)            // "aspace:cap/vote"
public class FleetVotes implements CapabilityProvider {

    private final VoteCapability delegate;

    public FleetVotes(AgentSpaces spaces, PeerIdentity identity) {
        this.delegate = new VoteCapability(
                spaces.group("orders-fleet").space("votes"),
                identity.agent("fleet-votes"), identity.peerId(),
                InstantSource.system());
    }

    @Override public String capabilityType() { return delegate.capabilityType(); }

    @Override public CapabilityAdvertisement describe(GroupId group) {
        return delegate.describe(group);
    }

    @Override public void start() { delegate.start(); }
}

You need this only when your peer is the one serving a capability. Most peers consume capabilities instead, which needs no annotation at all — see chapter 5. And with the starter you rarely write this either: agentspaces.capabilities.* already registers the aggregate, vote, semantic-discovery and key-wrap providers for you. Write the bean when you want a capability the starter leaves off (the ordered log, which needs a member set the starter cannot guess), when you want to configure one differently (a gossip learner with a real training step), or when you are publishing a capability of your own.

Where it livesA class implementing CapabilityProvider, registered by spaces.bind(bean) or by being a Spring bean. It is not an agent, so by itself it gets no card, no worker loop, and no @SpaceRef injection — a capability is constructed with the space and pipes it needs, explicitly, as above. If you also put @AgentSpec or a bound method on the class, it becomes both: the facade registers the capability and binds the agent, and then @SpaceRef fields are injected too. Mixing is legal but usually less clear than two classes.
Keep in mindThe annotation's value is documentation for discovery tooling; the provider's own capabilityType() is what actually governs registration, so a mismatch between them is confusing rather than fatal. The runtime calls start() once before the first advertisement and refreshes the advertisement on every card refresh, but it never calls anything periodically after that — capabilities that gossip need you to drive them, which chapter 5 covers in detail.
attributedefaultmeaning
value— (required)the capability type URI, e.g. aspace:cap/vote; informational for discovery tooling, the provider's own capabilityType() governs registration
group"" (all)the group to advertise into; empty means every registered group that has a capability runtime

Conventions they share

  • Space inference. Every space attribute is optional. Empty resolves to the group's sole registered space; several spaces without a name fail fast at bind time, listing the candidates.
  • Group selection. Every binding annotation takes an optional group, by application-facing name or GroupId value. A binder for a different group treats the member as not its own and skips it, so one bean can bind methods into several groups at once; the facade's bind(...) routes the bean to every group its annotations name.
  • Durations. Every duration accepts Spring-style simple strings ("10m", "500ms", "30s", "2h", "1d") and ISO-8601 ("PT10M").
  • Method shape. Bound methods take exactly one parameter: the entry, or for the vote annotations the Proposal (@Ballot) and the Decision (@OnDecision). @BidFunction returns double, @Ballot the option as a String (null abstains); the others return the next entry, a Contribution to an aggregate epoch, or void.
  • Names. An agent's name comes from @AgentSpec (or the class); bind(agent, name) overrides it, so several instances of one class can be several agents — a council of seats, one class.
  • Cards for free. Consumed and produced types are collected from the signatures and published on the agent's card, TTL-leased and auto-refreshed.
  • Nothing is reserved. The binder reads the annotations and ignores everything else on the class. Other frameworks' annotations, other interfaces, other methods and fields all keep working, which is why an Embabel @Action becomes a fleet worker by adding one annotation rather than by being rewritten.

Choosing between them

Three questions settle almost every case. Should exactly one agent handle this entry? If yes, @SpaceTake; if every interested agent should see it, @SpaceNotify. Are the candidate workers interchangeable? If not, put the space on AUCTION and add a @BidFunction, and let the cheapest one take it. Does one trigger produce more than one entry, or entries elsewhere? Then add @SpaceRef and write them yourself rather than contorting a return value.

The two failure models are the thing to hold on to, because they decide where risky work belongs. @SpaceTake retries forever by default: a crash, a kill, or a thrown exception all lapse the lease and the entry comes back. @SpaceNotify does not retry at all: a reaction that throws is logged and gone. Put work that must eventually happen behind a take; put reactions that are nice to have behind a notify.

aspace://guide/spring-boot · chapter 2

The Spring Boot starter

The starter (agentspaces-spring-boot-starter) wires the peer from properties and enrolls annotated beans through an ordinary bean post-processor. It deliberately declares no Spring dependency of its own; it compiles against a provided-scope stub surface and binds to the Spring your application already runs.

Properties

propertydefaultmeaning
agentspaces.keystore(ephemeral)directory persisting the peer's Ed25519 identity; omit for a fresh identity per run
agentspaces.bind(dial-only)host:port to listen on; omit for a NAT-restricted, dial-only peer
agentspaces.roles[]topology roles this peer volunteers for, e.g. RENDEZVOUS, RELAY
agentspaces.tick-millis250protocol tick period (membership probe, ad refresh, anti-entropy)
agentspaces.card-refresh-millis300000how often bound agents' cards republish
agentspaces.groups[].name—application-facing group name
agentspaces.groups[].founding—founding seed string; every member derives the same self-certifying GroupId from it. Exactly one of founding and join identifies the group; neither set means the group name doubles as the founding string
agentspaces.groups[].join—join an existing self-certifying group by aspace://<groupId> (or bare GroupId): the signed founding advertisement is fetched from the seeds and verified against the id before joining
agentspaces.groups[].join-timeout-millis10000how long the join-by-GroupId fetch waits for a verified answer
agentspaces.groups[].seeds[]existing members to dial, host:port
agentspaces.groups[].content-key—the group content key for payload encryption; also what the key-wrap capability serves
agentspaces.groups[].spaces[].name—a replicated space to attach
agentspaces.groups[].spaces[].strategyLEASE_RACELEASE_RACE or AUCTION; ORDERED is not a space strategy but a coordinator over the ordered-log capability (chapter 4)
agentspaces.groups[].spaces[].settle-window-millis200the take settle window
agentspaces.groups[].spaces[].admissiongroupgroup, allowlist, credential, or authorizer (chapter 4)
agentspaces.groups[].spaces[].allowed-agents[]the admitted agents under admission: allowlist
agentspaces.groups[].spaces[].credential-issuerthis nodeunder admission: credential, the PeerID whose signed credentials admit agents; only that node may grant and revoke
agentspaces.security.profilemtlsdev-local, mtls, mtls-oidc, zero-trust; selects transports, channel mode, and who answers authorization
agentspaces.identity.agent-keyspeerpeer signs every bound agent's records with the peer key; subordinate gives each bound agent a peer-certified Ed25519 key of its own, so its records are AGENT_ATTESTED, its card carries the key, and its @SpaceRef handles are per-agent views (chapter 6)
agentspaces.security.oidc.*—issuer, audience, jwks-url (all required under the OIDC profiles), plus this node's own token or token-file
agentspaces.security.grants.*(any member)under the membership profiles, PeerID lists narrowing one operation: raft-voter, directive-issuer, connector-serve, key-holder, space-write, space-take, vote
agentspaces.transport.channel-auth(from profile)signed signs every frame's envelope; attested negotiates unsigned frames over channels the transport authenticated
agentspaces.transport.tls.enabled(from profile)switch the TCP listener and dialer to TLS 1.3 with this peer's identity-endorsed channel certificate
agentspaces.transport.tls.trust-store—setting it switches to enterprise-CA mode: peers attest only when their chain validates to these anchors
agentspaces.transport.tls.key-store—this node's CA-issued credential (PKCS#12/JKS), with key-store-password / trust-store-password
agentspaces.transport.tls.crls[]cached CRL documents; revocation-strict fetches OCSP and distribution points live instead
agentspaces.transport.tls.require-attestationfalserefuse every frame not arriving on a channel attested for its sender, which is how CA revocation becomes a fleet-wide eject. Registers TLS alone, and refuses to start with TLS off
agentspaces.crypto.signature-providerjdkthe Ed25519 implementation by SignatureProvider name; every provider is pinned to the golden signature vectors
agentspaces.multicast.enabledfalseLAN bootstrap beacon: announce endpoints on multicast.group (default 239.255.42.99:7787) every interval-millis (5000) and dial peers heard, so a LAN fleet needs no seeds
agentspaces.capabilities.enabledtruethe capability runtime as a whole (chapter 5)
agentspaces.capabilities.*see noteper-capability toggles: aggregate, vote, semantic-discovery, key-wrap default on; ordered-log and gossip-learn default off
agentspaces.capabilities.votes-spacevotesthe space the vote capability runs over; when the group configures no space by that name the starter builds a private one, which does not count for single-space inference
agentspaces.console.enabledfalseserve the fleet console (chapter 7)
agentspaces.console.port7590console HTTP port
agentspaces.console.fleet-name""the fleet name the console displays
agentspaces.console.command.enabledfalsecommand-and-control routes; requires a token or a configured identity provider
agentspaces.console.command.token""static operator secret gating the command routes; when set it always wins over the identity provider
agentspaces.console.command.required-scopeaspace:console:operatewith the token blank and security.oidc.* set, the scope each operator's bearer JWT must carry
agentspaces.console.command.control-spacefleet-controlthe space directives are written to; workers subscribe here
agentspaces.console.command.audit-spacec2-auditthe durable command audit log
agentspaces.embabel.remote-actions-enabledtrueexpose discovered cards to the Embabel planner (chapter 3)
agentspaces.embabel.remote-action-timeout-millis30000how long a generated planner action waits for its result
agentspaces.embabel.auto-deploytrueredeploy the generated remote-action agent onto the platform whenever the set of remote actions changes
agentspaces.embabel.deploy-poll-millis2000how often the deployer checks the registry for arrivals and expiries
agentspaces.embabel.remote-agent-nameremoteFleetthe name of the generated agent

Security profiles and authorization

One property sets the whole posture. agentspaces.security.profile picks the transports and channel mode a fleet forms over — the default, mtls, means a fleet comes up over TLS with attested channels and no further configuration — and it also picks who answers the one question every privileged operation asks: authorizer.permits(peer, operation, scope).

Under dev-local and mtls the answerer is membership-rooted, one authorizer per group over the group's admitted members, narrowed by the optional grants lists; a grant never outlives admission. Under mtls-oidc and zero-trust it is your identity provider: each peer carries its JWT inside its own signed self-advertisement, the token's agentspaces_peer claim binds it to one PeerID, and scopes of the form aspace:<operation>[:<scope>] decide. Incomplete OIDC configuration fails at startup rather than at the first refusal.

# application.yml
agentspaces:
  security:
    profile: mtls-oidc
    oidc:
      issuer: https://idp.example.com
      audience: agentspaces-fleet
      jwks-url: https://idp.example.com/.well-known/jwks.json
      token-file: /var/run/secrets/aspace-token   # re-read each card refresh
  groups:
    - name: research-fleet
      spaces:
        - name: tasks
          admission: authorizer    # aspace:space-write:tasks / aspace:space-take:tasks
        - name: vault
          admission: credential   # issuer-signed SpaceCredential entries admit

The starter threads the authorizer into the key-wrap provider, spaces admitted by authorizer, the vote tally, and the DirectiveGates bean workers use to obey the console. Pass the Authorizer bean, or Authorizers.forGroup(name), to a RaftLog, DataQueryClient, or ConnectorRuntime you wire yourself.

What the starter does with your beans

After initialization, any bean declaring an agent identity or carrying bound methods enrolls automatically: bid functions wire first, worker loops start, subscriptions register, @SpaceRef fields inject, and the card publishes into every joined group. Group selection is by space coverage, narrowed by any group attribute the annotations carry: a bean binds into each configured group whose spaces cover the spaces it names, and a bean naming a space no group registers fails at startup, which is when a configuration mistake should surface. A @ProvidesCapability bean registers on the group's capability runtime the same way. Plain beans pass through untouched.

The facade bean is there when you outgrow the annotations: AgentSpaces navigates groups and spaces Embabel-style, with the typed verbs of chapter 4 as terminal operations.

// Ordinary application code that happens to need the fleet: inject the facade.
@RestController
public record OrderIntake(AgentSpaces spaces) {

    @PostMapping("/orders")
    public void submit(@RequestBody Order order) {
        spaces.group("orders-fleet").space("work")
              .write(order, Lease.of(Duration.ofMinutes(10)));
    }
}
Where it livesAnywhere Spring can inject a bean. A controller, a scheduled job, a message listener, a service — none of them need to be agents, and none of them get cards. The facade is also how you reach discovery (group.discovery()) and typed capability clients (group.capability(...)), which is why chapters 4 and 5 both start from it.
Keep in mindGroup selection happens at enrollment, from what a bean's annotations name, so a bean whose spaces no configured group covers fails at startup rather than quietly never running — which is the behaviour you want from a typo. The corollary: adding a space to an annotation means adding it to the YAML too.
aspace://guide/embabel · chapter 3

Embabel integration

The embabel-agentspaces extension activates only when Embabel is on the classpath, and it works in both directions: your Embabel agents become discoverable to the fleet, and the fleet's discovered skills become actions for your planner.

Outbound: cards from your @Agent beans

Every Embabel @Agent bean's metadata, its description, its @AchievesGoal statements, and its @Action input and output types, becomes a published AgentCard automatically. Worker loops still come from @SpaceTake on the action, so an Embabel action joins the fleet with one added annotation:

@Agent(description = "Researches topics from the shared task space")
public class ResearchWorker {

    @SpaceTake(space = "tasks", lease = "10m")
    @Action
    public Finding research(ResearchTask task, OperationContext ctx) {
        return new Finding(task.topic(),
            ctx.ai().withDefaultLlm().generateText("Summarize: " + task.topic()));
    }
}

Inbound: remote actions for the planner

The remote-actions bridge turns every foreign AgentCard in the group into a typed, invocable action. Invoking one is choreography, never RPC: the bridge writes the input entry where the remote workers take that type from, then awaits a result entry that correlates with the request (by default, equal values on every record component name the input and result share, the convention a Finding carrying its ResearchTask's topic already follows). A remote worker that crashes mid-step lets its lease lapse; whoever finishes, the awaited result still arrives. An invocation nobody serves times out empty with the entry still leased in the space.

RemoteActions fleet = RemoteActions.over(spaces.group("research-fleet"), peerId)
        .route(ResearchTask.class, "tasks")      // one-space groups need no routes
        .resultsIn(Finding.class, "tasks");

Finding f = fleet.producing(Finding.class).get(0)
        .invoke(new ResearchTask("agentic memory", 3),
                Finding.class, Duration.ofSeconds(30))
        .orElseThrow();

For the GOAP planner, EmbabelRemoteActions generates a class annotated @Agent with one real typed @Action method per remote action (card goals surface as @AchievesGoal), because the planner conditions on method signatures. The proven consequence: an agent whose local actions cannot chain to a goal type finds no plan alone, and finds and executes the multi-hop plan through remote specialists the moment the fleet's cards are bridged in. Cards are TTL-leased, so the planner's action space tracks the live fleet.

Where it livesWired once at startup, not on an agent: build the RemoteActions registry as a @Bean (or in assembly code), declare the routes, and hand it to EmbabelRemoteActions. The partybus does exactly this in one method, looping over its request and result types to register routes. Individual agents stay unaware — they are on the serving side of this, taking the entries the bridge writes.
Keep in mindCorrelation is by convention: the default matches equal values on every record component the input and result share, so a result type that carries none of its request's fields will never correlate and every invocation will time out. Either keep an id field on both, or supply your own Correlation. And because invocation is choreography, not RPC, a timeout means "nobody produced a correlated result in time" — the input entry is still sitting in the space under its lease, and a worker that shows up late will still do the work.

Deployment keeps itself current. With agentspaces.embabel.auto-deploy on (the default) a deployer watches the bridge's registry on a virtual thread and redeploys a freshly generated agent whenever the set of actions changes — the first time cards arrive, when a new card brings a new capability, and when an expired card takes one away. A card refresh that changes nothing but the issue time does not redeploy. Embabel's platform exposes no undeploy in the surface this module reaches, so an action that disappears is absent from the next generated agent while the previous generation stays registered until the platform restarts; invoking one of its vanished actions times out and throws, which the planner treats as the action failing.

Sizing the timeoutA generated action that gets no correlated result within remote-action-timeout-millis throws, which the planner treats as the action failing. Set it above your workers' realistic completion time, not above your patience.
aspace://guide/space · chapter 4

The Space API

One layer down from the annotations sits the interface they compile to: a typed, leased tuple space. Entries are plain records; every entry is signed by its writer and expires unless renewed; every peer holds a replica that keeps answering through partitions.

Getting a handle

Before the verbs, the practical question: where does a Space come from? There are three sources, and they return the same object — a space is a space however you reached it.

// 1. Inside an agent: @SpaceRef, injected at bind time. The common case.
@SpaceRef("findings") Space findings;

// 2. Anywhere else under Spring: the AgentSpaces facade bean, injected normally.
@Component
public record Coordinator(AgentSpaces spaces) {
    public void submit(String topic) {
        spaces.group("research-fleet").space("tasks")
              .write(new ResearchTask(topic, 3), Lease.of(Duration.ofMinutes(30)));
    }
}

// 3. Building one yourself: plain Java, tests, or a peer with no Spring at all.
ReplicatedSpace tasks = ReplicatedSpace.builder(runtime, "tasks", identity, "my-agent")
        .settleWindow(Duration.ofMillis(200))
        .build();
LocalSpace scratch = LocalSpace.builder("scratch", identity.agent("my-agent")).build();

// 4. Another agent's view of the same replica: same space, different signer and issuer (chapter 6).
Space asAuditor = tasks.as(identity.subordinate("auditor"));
Where it livesA controller, a scheduled method, a message listener, a test, a main — nothing about the space API requires an agent. The annotations of chapter 1 are a convenience over this interface, not a gate in front of it. Use the facade whenever the caller is ordinary application code that happens to need the fleet, and @SpaceRef when the caller is already a bound agent.
Keep in mindLocalSpace is the single-process form: same interface, same leases, same conflict behavior, no network and no peer. It is what you want in unit tests, and it is genuinely useful in production for in-process fan-out. Swapping it for a ReplicatedSpace changes no calling code, which is the point — write the logic against Space, test it locally, deploy it replicated. One replica per space name per node: a second ReplicatedSpace for a name that is already open is refused, naming the stream, and close() is what deregisters the first so the name can be opened again — example-02-research-fleet closes and reopens one.

The verbs

<T> EntryHandle          write(T entry, Lease lease);
<T> EntryHandle          write(T entry, Lease lease, Map<String,String> tags);
<T> Optional<T>          read(Template<T> template);
<T> Optional<T>          read(Template<T> template, Duration timeout);
<T> List<T>              readAll(Template<T> template, int limit);
<T> List<Issued<T>>      readAllIssued(Template<T> template, int limit);
<T> Optional<TakenEntry<T>> take(Template<T> template, Lease takeLease, Duration timeout);
    void                 complete(TakenEntry<?> taken);
<R> EntryHandle          complete(TakenEntry<?> taken, R result, Lease resultLease);
<T> Subscription         notify(Template<T> template, SpaceListener<T> listener, Lease lease);

take is destructive and exclusive under the space's conflict strategy; the taker holds a TAKE lease and must complete before it lapses, or the entry reappears. complete with a result validates and prepares the result record before committing the completion, so a task can never end up completed but resultless. Reads run against the local replica and never block on the network.

readAllIssued returns each match as an Issued<T>(T entry, AgentId issuer, Attestation attestation), where the issuer is authenticated by the entry record's signature and the attestation says how: PEER_ASSERTED when the hosting peer's key signed and vouched for the name, AGENT_ATTESTED when the agent's own peer-certified key signed (chapter 6). Any reader that makes a trust decision per entry — a vote tally, an audit reader, a result consumer — uses this rather than a field the entry declares about itself. SpaceEvent carries the same authenticated issuer.

notify events arrive at least once with entry deduplication available to you via the event's entry id; listeners must return promptly (the @SpaceNotify binder does both disciplines for you, which is why it is the recommended surface). Event kinds are WRITTEN, TAKEN, COMPLETED, EXPIRED (a write lease lapsed and the entry vanished) and REAPPEARED (a TAKE lease lapsed and the entry came back).

ReplicatedSpace adds a consistency hint on read and readAll. The default, ConsistencyHint.LOCAL, matches the replica as it stands; FRESH triggers exactly one anti-entropy pull toward one random member before matching locally — the cheap "make sure I am current" a late joiner or an operator console wants, without changing the local-first default for everyone else.

Templates and matchers

Template.of(ResearchTask.class)
        .where("kind", eq("summarize"))
        .where("priority", gte(3));

Matching is by type plus field predicates. The matchers (ai.badmonkey.agentspaces.api.space.Matchers): eq, ne, gt, gte, lt, lte, in(values...), contains(fragment) for character sequences, isNull, notNull, and predicate(fn) as the escape hatch. Fields resolve through record components; on a non-record type the usual getX/isX accessors are read instead. An unknown field name throws at where(...) time, listing the fields the type does have.

Leases and handles

Lease.of(Duration) is the only constructor you need. A write returns an EntryHandle with renew(extension) and cancel(); cancellation is issuer-signed, so nobody else can withdraw your entry. Everything shares the one failure idiom: state you stop renewing goes away, and nothing needs a cleanup job.

Conflict strategies: pick per space

The first two are strategies the space itself decides in its claim lattice, chosen when the space is built. The third is not a space setting at all — see the note below the table.

strategyhow a take is decidedcostreach for it when
LEASE_RACEfirst claim wins after one settle window (default 200 ms); at most one completion per entry, everone settle window of latency; transient duplicate work possible across partitionsthe everyday default; idempotent tasks
AUCTIONevery take carries a bid from the agent's bid function; the lowest bid wins, converging over gossipa few gossip rounds; needs a bid functioncost-aware routing: cheap capacity sweeps easy work, premium wins hard work
ORDEREDsigned take claims are submitted to a Raft log as commands; its commit order arbitrates identically on every member and the first valid committed claim per entry generation wins — exactly once, in one total ordermajority quorum for liveness: a partitioned minority cannot take at all, which is precisely the trade the strategy sellsnon-idempotent, high-stakes entries: payments, filings

A space fixes one strategy for all its entries; use several spaces the way you would use several queues (tasks-bulk, tasks-auction, commits).

What the claim stamp leaves openUnder LEASE_RACE and AUCTION a claim is judged partly against what the replica has already merged, so the outcome depends on arrival order. A back-dated first claim on an unclaimed entry wins at every replica it reaches before an honest one; two replicas that merged different first claims keep different winners, and anti-entropy does not reconcile them; and an honest claim arriving from a healed partition is refused where a newer claim merged first. Narrow TAKE scope, or reach for ORDERED, where strict fairness or convergence on the winner matters.
ORDERED is a coordinator, not a space settingPassing ConflictStrategyType.ORDERED to the space builder throws, and so does strategy: ORDERED in YAML. Build the space with the default strategy and take through the ordered-log capability instead, as below. Reads, writes, and completions still go through the space as usual; only take moves.

Taking through the ordered log

Exactly-once is one annotation now: @OrderedTake (chapter 1) is @SpaceTake with the take routed through the space's ordered-log coordinator, so the loop, the resubmit after a leader election, and the complete-or-lapse are the binder's. What stays yours is the quorum: a fixed member set, one coordinator per space, built with OrderedTakes.over(...) and registered on the group. The exactly-once desk and the Party Bus booking clerks are the worked examples.

// examples/example-12-exactly-once-desk -- ExactlyOnceDesk.startClerk(...)
OrderedTakes ordered = OrderedTakes.over(new CapabilityPipes(base.runtime(), codec), codec,
        identity, name, members, InstantSource.system(), seed, base.payments());
base.group().ordered(PAYMENTS, ordered);   // registers the log on the peer tick; @OrderedTake can now bind

// ExactlyOnceDesk.Clerk: the exactly-once worker is one method
@OrderedTake(space = PAYMENTS, lease = "30s", pollTimeout = "2s", resultLease = "1h")
public PaymentReceipt confirm(PaymentOrder order) {
    return new PaymentReceipt(order.orderId(), order.payee(), order.cents(), peer.name(),
            peer.raft().commitIndex());
}

Reach for this only where a duplicate is genuinely unacceptable and you would otherwise reach for a distributed lock: charging a card, filing a return, allocating a serial number, sending a message a customer will see once. Everywhere else LEASE_RACE is cheaper and more available, and the duplicate work it can produce across a partition is the price of continuing to run at all.

Where it livesThe worker is an annotated method on any bound class; the quorum is assembly code that runs once. The RaftLog and OrderedTakes are wired at startup — a @Bean method, or the assembly code of a plain-Java peer — and the worker is a thread you own. If you want the rest of the agent model around it, keep the ordered worker as its own class and let the annotated agents handle everything that is not the exactly-once step, which is how the partybus keeps one small Raft quorum for bookings and leaves thirty-odd other agents on LEASE_RACE.
Keep in mindThe member set is fixed and every member must be given the same list: this is a static quorum, not a group that grows with membership. A minority partition cannot take at all, by design — that is the availability you are trading away for exactly-once. The circular wiring above (an AtomicReference so the log's apply callback can reach the coordinator that has not been constructed yet) is genuinely necessary; both halves need each other. Registering the log is the whole of the clock wiring: the peer tick drives elections and heartbeats, the log is told the real cadence so the leader lease it advertises is truthful, and because the runtime republishes an advertisement whose content changed on the very next tick, role=leader is discoverable within a quarter-second of an election rather than on the next five-minute card refresh. A log you construct but never register is never elected, and every take times out empty. Every replica — including one that holds no claim of its own, like the desk — authenticates a completion against the holder's own signed claim, which the log carries through unchanged; that is what lets finished work drain everywhere instead of reappearing when its TAKE lease lapses.

Space admission: who may write and take

A space advertises how it admits writers, independently of how the group admits members. Reads stay open to the group in every mode.

  • GROUP (the default) admits every admitted member of the group.
  • ALLOWLIST admits a fixed set of agents. A local caller outside the list gets a SpaceAdmissionException; an inbound record or claim from an unlisted agent is dropped and reported as misbehavior.
  • CREDENTIAL admits whoever holds a live credential from the issuer. A credential is an ordinary leased SpaceCredential entry the issuer writes into the space itself, so nothing new appears on the wire: the issuer's node grants with space.grant(agent, scopes, validity) and revokes with space.revoke(agent), which is just a signed cancellation. A transition whose credential has not merged yet is dropped without a strike, because the evidence can legitimately arrive after the write it authorizes.
  • AUTHORIZER hands the decision to the profile's Authorizer, scoped by the space name, which is how an identity provider's scopes (aspace:space-write:tasks) end up deciding.

Each transition is judged once as it arrives, after its signature verifies. A refused local call throws; a refused remote delta is dropped, and strikes its sender only when the rule says the refusal cannot be a late credential. The judged identity is the acting agent, never the peer: when several agents of one peer share a replica through views (space.as(identity), chapter 6), each view is admitted or refused on its own agent's standing, so an ALLOWLIST or CREDENTIAL space can admit one agent of a peer and refuse its sibling. A record whose spaceId names another space is dropped before any of this, however validly signed.

Building a space

The starter builds spaces from YAML; plain Java builds them directly. LocalSpace is the single-process form (same API, ideal for tests); ReplicatedSpace attaches to a joined group's runtime.

ReplicatedSpace tasks = ReplicatedSpace.builder(runtime, "tasks", identity, "my-agent")
        .settleWindow(Duration.ofMillis(200))
        .writer(identity.subordinate("my-agent"))   // sign as the agent's own certified key (chapter 6)
        .strategy(ConflictStrategyType.AUCTION)     // with .bidFunction(fn)
        .blocks(new BlockExchange(runtime, codec))   // content-addressed payloads
        .groupKey(key)                               // AES-256-GCM payload encryption
        .tagShard(tags -> "eu".equals(tags.get("region")))  // partial replica
        .admission(SpaceAdmission.allowlist(agents))  // or group/credentials/authorizer
        .maxLease(Duration.ofHours(24))              // ceiling on any lease written here
        .schemaHints(ResearchTask.class, Finding.class)  // published on the space ad
        .advertise(publisher)                        // with .advertisementTtl(ttl)
        .build();

A space publishes its own advertisement the moment it is created, carrying its strategy, admission, replication mode, and schema hints, and refreshes it on the group's tick. When its founders' lease lapses the space turns read-only rather than silently accepting writes nobody will keep.

Payloads over 64 KiBConfigure .blocks(...) and large entries travel content-addressed: the space replicates a hash reference and peers fetch the bytes on demand, verified against the hash. Because reads never block on the network, an entry whose block has not landed yet is simply not visible yet: the replica starts one background fetch as soon as the record merges (a read that misses it starts one too), a read with a timeout polls within that timeout and sees the entry once the block arrives, a notify subscriber receives the entry's WRITTEN event the moment the block is local without anyone having to read, readAll skips it for that round, and a subscription event fires on the later merge or read that finds the block locally. Nothing is lost — it arrives a moment late.
aspace://guide/discovery · chapter 5

Discovery and capabilities

Below the space sit the layers that answer "who can do what" and "what does the fleet think": signed, TTL-leased advertisements, and capability services built over the same fabric.

Finding agents, assets, and services

Everything discoverable is a typed, signed card in a per-group cache that gossip keeps fresh; a card whose issuer stops republishing simply ages out. The advertisement families are peers, groups, spaces (with their admission and replication modes), capabilities, AgentCards (with their space bindings), AssetCards, and revocations. The cache verifies every signature, bounds its own occupancy, and evicts by TTL; a peer volunteering as RENDEZVOUS holds four times as much. Query structurally, or by meaning through the semantic-discovery capability (its reference embedder needs no model at all):

// who can turn a ResearchTask into a Finding?
List<AgentCard> researchers = discovery.find(AgentCard.class,
        card -> card.produces().contains(Finding.class.getName() + "#v1"));

// or escalate beyond the local cache: rendezvous peers first, then scoped gossip.
// The remote predicate is structural (field name to expected value) so it can
// travel on the wire, and it is validated against the type before anything is sent.
discovery.remoteFind(AgentCard.class,
        Map.of("agent", "researcher"), Duration.ofSeconds(2));

Most fleets never call this. Work finds its worker because the worker is sitting on the space taking that type, not because anybody looked anything up — the routing is the entry's type. Discovery earns its keep when you need to know who is out there rather than get one task done: a console drawing the room, a planner deciding whether a capability exists before committing to a plan (chapter 3 builds on exactly this), an operator asking which peers can still do a thing, or a dispatcher that wants to fail loudly when nobody is listening instead of writing an entry into a space where it will sit until its lease lapses.

Where it livesAnywhere you can reach a DiscoveryService: spaces.group("g").discovery() under Spring, or the service you constructed on a plain-Java peer. Not an agent concern — agents publish cards automatically by being bound, and read them only when they want to reason about the fleet.
Keep in mindfind is a local, in-memory scan of the cache, so call it as often as you like. remoteFind puts a query on the wire and blocks for the wait you give it, so it belongs on a slow path, not in a loop. Both return what the cache holds now: a card whose issuer died is still there until its TTL runs out, and a card from a peer that just joined may not have arrived yet. Treat the result as a hint about a live fleet, never as an authoritative roster — which is the same discipline the fabric itself applies.

How a capability is used

Layer 4 is the part of the system people most often wire wrongly, because a capability is not one thing but three, and a given peer may play any subset of the roles.

  • The provider is the object that implements CapabilityProvider and serves the capability on this peer. Under the starter, the shipped ones are registered for you from agentspaces.capabilities.*. Otherwise it is a @ProvidesCapability bean (chapter 1).
  • The consumer is whatever calls it. Typed clients resolve off the group context — spaces.group("g").capability(VoteClient.class) — and need no annotation at all. Most application code is only ever this.
  • The clock is what advances a continuous capability's own protocol — a push-sum exchange, a gossip-learning offer, a Raft election step. You do not supply it. A capability runtime registers its protocol tick with the group's peer tick when it is constructed, so a provider is driven because it is registered, the same way a replicated space converges because constructing it joins that same clock. Nothing to schedule, nothing to remember.

That last point is the one worth internalising, because it is what makes the annotations of chapter 1 enough on their own. One clock runs the whole stack: membership probing, gossip, space anti-entropy, and capability protocols all advance on the peer's tick, at agentspaces.tick-millis (250 ms by default). Set that one property and you have set the convergence rate of the entire fabric.

// There is no driver to write. Registering the capability is the whole of it:
// under the starter this already happened, and under plain Java it happens
// the moment the runtime exists.
AgentSpaces.GroupContext group = spaces.group("orders-fleet");
group.capabilities().isDrivenByPeerTick();     // true

// A test that wants each round to be an exact, countable step takes the clock
// back and drives it by hand -- the same escape hatch a TestClock is for.
group.capabilities()
        .detachFromPeerTick()
        .cadence(Duration.ofMillis(300));      // what you will really tick at

group.capabilities().protocolTick();           // one round, deterministically

Convergence takes a few dozen rounds, so the peer tick period is what turns "converges in O(log N) rounds" into a wall-clock number: at the default 250 ms a fleet settles in a few seconds. Vote, semantic discovery and key wrap are request/response and need no clock at all, which they declare through requiresTick(); the runtime drives them anyway, because the default tick() is a no-op and costs nothing.

Two clocks exist at this layer and they are deliberately different rates. The protocol clock above runs at the peer tick. The advertisement clock, refreshTick(), republishes leased capability advertisements and runs with the card refresh, on the order of minutes. A method named for one will not do the other.

The six capability services

  • Vote (aspace:cap/vote): proposals and ballots as signed entries; the decision closes at quorum and any member recomputes the tally from the space. QUORUM mode is the canonical, auditable form, counting one ballot per eligible voter against an electorate of fresh AgentCards narrowed by the authorizer; MAJORITY_GOSSIP is a cheap approximate temperature check by push-sum. The voter is derived from the entry's authenticated issuer, never from a field the ballot declares.
  • Aggregate (aspace:cap/aggregate): push-sum gossip; every member converges on a fleet-wide figure with no collector. Six operators — SUM, AVG, COUNT, MIN, MAX, QUANTILE — over a supplied value or over a template's matching entries. aggregate.start("backlog", localValue), then estimate("backlog").
  • Ordered log (aspace:cap/ordered-log): a small Raft among volunteers, with the leader lease published as an advertisement; backs the ORDERED strategy and any workflow needing a total order. Off by default, since it needs a fixed member set.
  • Gossip learning (aspace:cap/gossip-learn): mergeable models exchanged pairwise converge to the fleet mean with no parameter server; content-addressed, with a mass-conserving offer-and-accept exchange. Suits learned bid policies. Off by default; enabling it wires the default weight-averaging learner.
  • Semantic discovery (aspace:cap/semantic-discovery): an embedding index over the group's advertisements behind a pluggable Embedder. query(text, limit) searches locally, remoteQuery(text, limit, timeout) escalates; the shipped hashing embedder needs no model at all.
  • Key wrap (aspace:cap/key-wrap): distributes group content keys sealed per member (X25519 + HKDF + AES-256-GCM), so encrypted spaces onboard members without any shared secret in transit.

Typed capability clients

Consumers do not speak the capability's entries directly. A client type resolves off the group context, and the client finds the providers, speaks the protocol, and hands back ordinary values:

VoteClient vote = spaces.group("orders-fleet").capability(VoteClient.class);
vote.propose("rollout-7", "ship the new pricing model?",
            List.of("yes", "no"), 3, Lease.of(Duration.ofMinutes(10)))
     .castBallot("rollout-7", "yes", Lease.of(Duration.ofMinutes(10)));

Optional<VoteCapability.Decision> decided = vote.decision("rollout-7");  // or tally(id)

AggregateClient agg = spaces.group("orders-fleet")
        .capability(AggregateClient.class);
agg.start("backlog", queueDepth()).estimate("backlog");

Register your own with spaces.clientFactory(new MyClient.Factory()), and serve your own capability with a @ProvidesCapability bean (chapter 1).

Who counts as a voterA ballot names its voter, and the tally discards any ballot whose declared voter is not the entry's authenticated writer — that check is what stops anyone voting on your behalf. How many distinct voters a peer can be is then the authorizer's decision, not the ballot's: the tally counts at the authorizer's granularity(VOTE, space). Under PEER (the default, and every OIDC profile) one peer is one counted ballot however many agent names it writes, so a member cannot multiply itself. Under AGENT — switched on by naming peer/localName AgentIds in agentspaces.security.grants.vote — each named agent counts once, and only when its ballot is AGENT_ATTESTED, signed by the agent's own certified key (chapter 6). A fresh AgentCard is a liveness filter, never an eligibility check. Decision.granularity() says which rule closed the vote. A VoteCapability still votes as the one agent its space handle writes as (space.writer()), and constructing one with another identity throws rather than handing you an object whose every ballot would be silently dropped.

Reaching a capability from inside an agent

An agent POJO has @SpaceRef for spaces; the equivalent for capabilities is @CapabilityRef, and for what agents do with a capability there are annotations of the same shape as @SpaceNotify: @Ballot and @OnDecision for the vote, @OrderedTake for the ordered log, a Contribution return for the aggregate (chapter 1). Constructor injection of the capability object still works and is still fine when the caller has it in hand; the annotations are the idiomatic form because the discipline they encode — one ballot per proposal, one reaction per closed vote, the resubmit after an election — is what every hand-written consumer got wrong at least once.

// agentspaces-partybus -- Party.TravelerAgent: contributes to the aggregate and votes, with no vote code
@AgentSpec(name = "traveler", description = "Reports energy each day and votes on the party's choices")
public static final class TravelerAgent {
    private final PushSumAggregate aggregate;      // constructor-injected, or @CapabilityRef AggregateClient
    @SpaceRef(TRIP) Space trip;

    @SpaceNotify(space = TRIP, lease = "12h")
    public EnergyReport feel(TripDay day) {
        int energy = energyOn(day, ...);
        aggregate.start(Itinerary.energyEpoch(day.tripId(), day.date()), energy);
        return new EnergyReport(day.tripId(), day.date(), name(), energy);
    }

    @Ballot(space = DECISIONS, prefix = "next:", lease = "12h")
    public String decide(VoteCapability.Proposal proposal) {
        return preference(proposal.options());       // the cue is the proposal itself: no wait, no voted set
    }
}

Under Spring the typed clients resolve from the facade — spaces.group("trip").capability(VoteClient.class) — or arrive by @CapabilityRef; SemanticClient joins VoteClient and AggregateClient. Providing a capability through the facade (group.provide(vote)) is what makes the vote annotations on its space bindable; group.ordered(space, coordinator) does the same for @OrderedTake.

Where it livesCapabilities are peer-scoped, not agent-scoped: one PushSumAggregate per peer serves every agent on it, and one vote capability per vote space casts for every bound agent — as that agent, when the agent has a key of its own. There is no driver to get wrong: the capability runtime drives what it registers on the peer tick.

Worked example: sensing the fleet with aggregate

Push-sum answers questions no single peer can answer: what is the fleet's average backlog, how many workers are live, what is the p95 of a number every peer holds a piece of. There is no collector and no metrics server; every member ends up holding the same estimate.

// Every member contributes its own number to a named epoch...
aggregate.start("backlog", myQueueDepth());

// ...the peer tick exchanges them, and after a few dozen rounds every member
// reads the same fleet-wide figure. Empty until the first round completes:
// a capability that has not run yet will not pass its own seed off as a
// fleet answer.
OptionalDouble sensed = aggregate.estimate("backlog");

// A supervisor can then act on what the whole fleet feels, not on its own view:
if (sensed.orElse(0) > BACKLOG_THRESHOLD) {
    vote.propose("scale-up-1", "Average backlog is " + sensed.getAsDouble()
            + "; add two workers?", List.of("approve", "reject"),
            4, Lease.of(Duration.ofMinutes(10)));
}

That pair — sense with aggregate, decide with vote — is the whole of examples/example-05-quorum, and it is the canonical shape for autonomic behavior: the fleet notices its own condition and commits to a response that every member can audit afterwards from the vote space.

Keep in mindAn epoch id is just a string every participant agrees on; different strings are different, independent aggregations, which is how the partybus runs one epoch per trip day. estimate returns empty until this peer has joined the epoch and completed at least one exchange round, and returns a converging number before it has settled — read it a moment after start, not in the next statement. An estimate that stays empty is the fabric telling you the capability is not being driven, which under the starter should never happen. Over a fully replicated space, prefer AVG, MIN, MAX and QUANTILE in template mode: SUM and COUNT over entries every replica already holds will count each entry once per replica, which is almost never what you meant.

Worked example: gossip learning

This is the least obvious capability, so here it is end to end (and as a runnable fleet: examples/example-14-gossip-learning, three peers whose local weight vectors converge on the fleet mean on the peer tick alone, with a test that asserts it). The idea: every peer holds a parameter vector, peers pair off at random and average their vectors, and the fleet converges on the mean without a parameter server and without any raw data leaving any peer. Add a local update step that runs after each merge and you have federated training — train on what you have, average with a neighbour, repeat.

Two things have to be true for it to work, and both are your job: every participant registers the same model id with the same vector length, and the local update (if any) is supplied at construction. The third — something driving the exchange — is the runtime's job, and it is done for you as soon as the provider is registered.

// 1. Build the learner. The starter's gossip-learn toggle gives you a pure
//    averaging learner; to train, construct your own and publish it as a bean.
@Component
@ProvidesCapability(GossipLearning.TYPE)          // "aspace:cap/gossip-learn"
public class BidPolicyLearner implements CapabilityProvider {

    private final GossipLearning learning;

    public BidPolicyLearner(AgentSpaces spaces, CborCodec codec) {
        AgentSpaces.GroupContext group = spaces.group("orders-fleet");
        // The builder is the configurable form: a training step, and a metrics
        // space so endEpoch below has somewhere to write.
        this.learning = new GossipLearning(GossipLearner.builder(
                        group.pipes(), group.runtime().sampler(),
                        group.identity().peerId(), codec, group.clock(),
                        WeightAveraging.INSTANCE)
                .localUpdate(params -> sgdStep(params, batch()))  // the training step
                .metrics(group.space("metrics"), Lease.of(Duration.ofDays(7)))
                .build());
    }

    @Override public String capabilityType() { return learning.capabilityType(); }
    @Override public CapabilityAdvertisement describe(GroupId g) {
        return learning.describe(g);
    }

    // 2. Join the model. Every peer must use the same id and the same length.
    @EventListener(ApplicationReadyEvent.class)
    public void join() {
        learning.start("bid-policy", new double[]{1.0, 0.0, 0.5});
    }

    // 3. Advance the protocol. Inherited from CapabilityProvider and called by
    //    the peer tick because this bean is registered -- you write no driver.
    @Override
    public void tick() {
        learning.tick();
    }

    // 4. Read the current parameters wherever you need them -- e.g. from a
    //    @BidFunction, which is what makes the fleet's bid policy learned.
    public double[] parameters() {
        return learning.model("bid-policy").orElse(DEFAULTS);
    }
}
// And at the end of a training epoch, write one evaluation per model into the
// metrics space, so the fleet's loss curve is itself replicated state:
List<GossipLearner.Evaluation> written =
        learning.endEpoch(epoch, params -> lossOn(params, holdoutSet()));

// Any member can then read everyone's, attributed to its authenticated author:
List<Space.Issued<GossipLearner.Evaluation>> all = metrics.readAllIssued(
        Template.of(GossipLearner.Evaluation.class).where("epoch", eq(epoch)), 100);

Use it for anything where peers hold private data and want a shared model: learned bid policies (the fleet converges on what work actually costs), drift detectors, shared bandit statistics, federated fine-tuning signals. It is the natural partner of @BidFunction — a bid function is a plain method over numbers, and this is how those numbers learn.

The mechanism is randomized pairwise averaging, which contracts the spread geometrically, so a fleet lands within tolerance of the mean in O(log N) rounds. Every exchange conserves mass: the fleet mean is exactly preserved by each pairwise average, which is why the answer does not drift as peers join and leave mid-run.

Where it livesA provider bean, not an agent — and the four numbered steps above deliberately sit in one class because they share the learner instance. Consumers (a bid function, a scheduled reporter) read parameters() off that bean by ordinary injection. For a model that is not a double[], implement MergeableModel and build a GossipLearner<YourType> directly; GossipLearning is just the double[] specialization under WeightAveraging. Its two short constructors cover pure averaging and averaging-plus-update; anything more — a metrics space, a block exchange — goes through the builder as above.
Keep in mindPeers that never called start for a model id ignore exchanges about it, and a vector of the wrong length is dropped rather than merged — so a rolling deploy that changes the model's dimension silently partitions your fleet into two non-interacting populations. Version the model id when the shape changes. The local update runs after every merge, so keep it to one cheap step; a full training run in there will stall the tick. Models above the 64 KiB inline limit travel content-addressed through the block exchange automatically, but only if the learner was built with one. And endEpoch throws unless the learner was given a metrics space, which is the one reason to prefer the builder over the plain constructors.

Semantic discovery and key wrap

The two request/response capabilities need no driver and no model set-up. Semantic discovery answers "who holds something like this?" against the group's advertisements; key wrap hands a new member the group content key without ever putting it on the wire in the clear.

// Semantic: find an asset by meaning rather than by name (examples/example-10)
var matches = semantic.remoteQuery("who holds customer order history?", 3,
        Duration.ofSeconds(2));
AssetCard card = (AssetCard) matches.get(0).advertisement();

// Key wrap: the holder serves the key under the profile's authorizer...
keys.serve(contentKey, authorizer, groupId);

// ...and a joining member asks the holder for it, sealed to its own X25519 key.
Optional<GroupKey> key = keys.request(holderPeerId, Duration.ofSeconds(2));
Keep in mindSemantic matching is only as good as the embedder: the shipped HashingEmbedder needs no model and is a lexical approximation, good enough to find an asset by description and not good enough to be called understanding. Swap in a real Embedder when it matters. For key wrap, the authorizer is consulted per request, not once at serve time, so revoking a grant takes effect on the very next ask — and a request nobody serves completes empty rather than hanging.
aspace://guide/peering · chapter 6

Peering, identity, and transports

The floor of the stack. Most applications configure it through the starter and never touch it; this is the API when you wire by hand or embed without Spring.

PeerIdentity identity = PeerIdentity.generate();      // or FileKeystore.loadOrCreate(path)
PeerNode node = PeerNode.builder(identity).build();
node.listen(new TcpTransport(), "0.0.0.0:7500");      // and/or new QuicTransport()

GroupRuntime runtime = node.joinGroup(groupAd,
        new GroupMembership.Config(Duration.ofSeconds(30),  // member TTL
                                   Duration.ofSeconds(2),   // ping timeout
                                   2),                       // indirect probes
        seeds);
node.startTicking(Duration.ofMillis(250));
  • Identity. An Ed25519 keypair; the PeerId is the hash of the public key, so identity survives every address and transport change, and every frame, card, entry, and state transition the peer emits is signed with it.
  • Groups. A group is self-certifying too: its GroupId is the hash of a founder-signed founding document, so every member derives the same id from the same founding string, and a newcomer who knows only the id can ask a seed for the document and verify the answer against the id before joining. That is the founding and join pair of chapter 2, and joinGroup has an overload for each.
  • Membership. SWIM-style probing with leased membership: a silent member drops from the view on its own. Groups admit by policy: OPEN, INVITE (founder-signed credential presented at join), or POLICY (your validator SPI). Two gossip channels carry state, rumor for speed and anti-entropy for completeness, the latter paced by the group's own gossip period.
  • Topology. Any member may volunteer as RENDEZVOUS (a bigger discovery cache newcomers ask first) or RELAY (forwards signed frames verbatim for NAT-restricted peers). Correctness never depends on either. On a LAN, the opt-in multicast beacon replaces seeds entirely.
  • Transports. An in-JVM loopback for tests, TCP, TLS 1.3, and QUIC (RFC 9000) ship; a peer advertises (transport, addr, priority) endpoints and dialers try them in ascending priority, skipping schemes no configured transport serves. Peer identity stays with the signed envelopes on every transport.
  • Channel authentication. The envelope signature authenticates one hop, so when the transport has already authenticated the channel it is repeated work at roughly a thousand times the cost of serializing the frame. TLS and QUIC present a channel certificate the peer identity endorses, and a node that announces itself willing accepts unsigned frames from the peer that channel attests — the attested mode of chapter 2. Signed and attested nodes mix freely; forcing signed mode is the only downgrade an attacker can cause, and it costs nothing but speed.
  • Enterprise CA mode. Point the TLS binding at a trust store and peers attest only when their chain validates to your anchors, with revocation checked against cached CRLs or fetched live. Under require-attestation every inbound frame must arrive on a channel attested for its sender, which is what makes the CA's revocation an authoritative fleet-wide eject rather than advice.
  • Revocation and rotation. A signed RevocationAdvertisement gossips like any other advertisement, with anti-entropy so late joiners converge, and a revoked identity is refused at dispatch ahead of every other check, evicted from membership, and its connection closed. Because a self-certifying identity cannot rotate its key in place, rotation is a revocation that names a successor; the successor joins and is authorized like any new peer. GroupRuntime.revoke(...) and rotate(...) issue, and the founder-rooted RevocationValidator is the seam an enterprise deployment replaces.
  • One clock. startTicking(period) is the whole scheduler. Membership probing, rumour and anti-entropy gossip, replicated-space convergence, and capability protocols all advance on that one tick, because each layer registers with it rather than bringing a timer: a space through gossip().reconcile(...), a capability runtime through GroupRuntime.onTick(...), which is also where you hang periodic work of your own. Work registered there runs on the tick thread, so it must be quick; work that throws is logged and cannot stop the tick or its siblings.
  • Containment. Per-issuer rate limits and bounded strike counters, fed only by witnessed violations on fresh authenticated frames, quarantine a misbehaving peer. Replayed frames and unadmitted peers deliberately never strike, which is what keeps quarantine unweaponizable.
Where it livesAssembly code that runs once: a main, a test fixture, or your own bootstrap class. Under the starter you write none of it — the properties of chapter 2 build exactly this and the PeerNode is a bean you can inject if you want to assert the posture you configured. Reach for the hand-wired form when embedding a peer in something that is not a Spring application, when writing tests against a simulated network, or when you need a topology the properties do not express.
Keep in mindNothing happens until startTicking: the tick drives membership probing, advertisement refresh and anti-entropy, so a peer that joined but never ticks will look joined and converge with nobody. Join order does not matter and seeds need not be reachable forever — they are a way in, not a dependency. Close the node when you are done; an abandoned peer stays in other members' views until its membership lease lapses.

Agent identity and certificates

An AgentId is <peerId>/<localName>, and by default the name is the peer's word: the peer key signs every record and a reader trusts that the peer labelled its agents honestly. When per-agent non-repudiation matters — an audit trail that must hold this agent, not its host, to a finding — give the agent a key of its own. The peer certifies it once, for a validity window, and from then on the records that agent writes are signed by the agent key and travel with the certificate. Readers see the difference: every Space.Issued carries an attestation, PEER_ASSERTED for the default and AGENT_ATTESTED when the agent's own key signed.

// examples/example-13-signed-agents -- SignedAgents.startAgent(...)
// The agent's own key, vouched for by the peer: the one line that opts in.
AgentIdentity writer = identity.subordinate(name, issued, ttl);
ReplicatedSpace findings = ReplicatedSpace.builder(runtime, FINDINGS, identity, name)
        .writer(writer)
        .build();

// SignedAgents.ledger(...) / describe(...): who said it, and how strongly the fleet can hold them to it
for (Space.Issued<Finding> issued : reader.findings().readAllIssued(Template.of(Finding.class), 100)) {
    System.out.println(issued.issuer().encoded() + " (" + issued.attestation() + ")");
}
// zAuditorPeer.../auditor (AGENT_ATTESTED)   zDeskPeer.../desk (PEER_ASSERTED)

// examples/example-13-signed-agents -- SignedAgents.startAnnotatedPeer(...): two POJOs, one peer, no identity code
@AgentSpec(name = "auditor", description = "Audits releases", goals = {"audit"})
public final class Auditor {
    @SpaceRef(FINDINGS) Space findings;                      // arrives as this agent's view
    public void conclude(String subject, String verdict) { findings.write(new Finding(subject, verdict), FINDING_LEASE); }
}
// The starter does exactly this from agentspaces.identity.agent-keys=subordinate.
AgentSpaces spaces = new AgentSpaces(identity, InstantSource.system(), identity::subordinate);
group.bind(new Auditor());  group.bind(new Clerk());   // two provable identities on one node
  • What the certificate is. AgentCertificate(agent, agentPublicKey, issued, ttl, peerSignature): the peer's Ed25519 signature over the other four fields in canonical CBOR. Only the peer an AgentId names can issue one, and it is valid until issued + ttl on the receiver's clock. PeerIdentity.subordinate(name) mints a fresh keypair and certifies it for an hour; the overloads take your own KeyPair (for persistence or an HSM) and your own validity window.
  • What a replica checks. The peer key still travels beside the record and must hash to the issuing peer, as it always has. With a certificate present, the replica verifies the certificate under that peer key, requires it to name exactly the record's issuer and to be unexpired, and then verifies the record under the certificate's agent key. A record that arrives with a certificate but a peer-key signature is refused, and so is one under a lapsed certificate — example-13's second flow test shows a stale agent's findings landing locally and nowhere else.
  • Several agents, one peer. space.as(identity) is a view of the same replica seen as another agent of the same peer: its writes are signed and attributed as that agent, its takes name that agent as holder, and admission is judged for that agent, so two views of one replica can be admitted differently. One replica, one gossip registration, one clock per node stay the rule. Under the annotations you never call it: with agentspaces.identity.agent-keys=subordinate (or new AgentSpaces(identity, clock, identity::subordinate)) every bound agent gets a certified key and every @SpaceRef arrives as its view.
  • What it does not do. A certificate proves who signed, never what they may do: admission and authorization are still the space's and the profile's decision. Take claims and completions still carry the peer's proof (the holder is named by AgentId), and subordinate keys take no part in key wrap.
  • Wire compatibility. The certificate rides in an appended field of the state delta that is omitted when absent, so a fleet that never opts in is byte-identical to before, and older peers ignore the field. Python and TypeScript peers verify certified records; the golden vectors pin the certificate, the agent-key record signature, and five hostile deltas all three implementations refuse identically.
Where it livesUnder the starter, one property: agentspaces.identity.agent-keys=subordinate, and the annotations of chapter 2 do the rest. By hand, the facade's third constructor argument (identity::subordinate) or, for a single space, .writer(identity.subordinate(name)) on the builder and space.as(identity) for further agents. You read it back through Space.Issued.attestation() and AgentCard.attested().
Keep in mindA certificate that lapses mid-flight stops the agent's new records from replicating while its own replica still shows them, which looks like a partition from the agent's side. Pick a TTL that outlives the process, or re-issue on a cadence; the peer key, not the agent key, is what membership and key wrap are rooted in, so a subordinate key is cheap to rotate.
aspace://guide/console · chapter 7

The fleet console

A read-only peer that serves the whole room: a single-file web UI at /, a HAL+JSON API under /api/v1 walkable from the root by any hypermedia client, and an SSE activity stream at /api/v1/events, all derived from the replicated coordination state, with zero instrumentation in any agent. It binds to loopback by default, as the A2A gateway does. Two properties turn it on (chapter 2); two more add command-and-control: dispatch tasks, cancel what the console dispatched, and pause, resume, or drain workers, every command executed under the console peer's signed identity, attributed to the named operator, and recorded in a durable audit space.

Operators authenticate one of two ways. A static command.token gates the HTTP command routes and always wins when set; leave it blank with security.oidc.* configured and each request's bearer JWT is validated against your identity provider instead, required to carry aspace:console:operate, and audited under the token's own subject rather than a name the request declares. Directives travel through the fleet-control space and the audit trail lands in c2-audit; both names are properties, and both spaces belong in the group's spaces list.

Extend the read side two ways: a ConsoleViewCustomizer bean tells the default generic view what the application's entries mean (worker attribution above all), and every ConsolePanel bean becomes one more section on the page and one more document under /api/v1/panels. Extend the command side with ConsoleCommand beans, whose typed fields render as forms. Workers honor directives through a DirectiveGate — under the starter, gates.attach(space, worker) off the DirectiveGates bean — consulting gate.paused() and gate.draining() in the work loop. A gate obeys only the configured console peer, narrowed further by the directive-issuer grant, and rejects future-stamped directives.

// Teach the generic view what your entries mean. Worker attribution above all:
// without this the console can show that work happened but not who did it.
@Bean
ConsoleViewCustomizer researchConsole() {
    return (builder, spaces) -> builder
            .results("findings", Finding.class, f -> ((Finding) f).worker());
}

// One panel bean = one section in the UI and one document at /api/v1/panels.
@Bean
ConsolePanel storyPanel(ConsoleView view) {
    return new ConsolePanel() {
        @Override public String id()    { return "story"; }
        @Override public String title() { return "How this console works"; }
        @Override public Object data()  {
            return Map.of("source", "replica state and leased advertisements only",
                          "eventsObserved", view.lastSeq());
        }
    };
}
// A worker that obeys the console. The gate is a subscription on the control
// space; consult it at the top of the loop, and close it when the worker stops.
try (DirectiveGate gate = gates.attach(controlSpace, "worker-1")) {
    while (!Thread.currentThread().isInterrupted()) {
        if (gate.draining()) {
            return;                       // commanded to drain: stop for good
        }
        if (gate.paused()) {
            Thread.sleep(200);            // commanded to pause: hold off taking
            continue;
        }
        Optional<TakenEntry<TaskEntry>> taken = tasks.take(Template.of(TaskEntry.class),
                Lease.of(Duration.ofSeconds(30)), Duration.ofSeconds(2));
        taken.ifPresent(t -> tasks.complete(t, work(t.entry()), RESULT_LEASE));
    }
}

Turn it on for any fleet somebody has to operate. It is a peer like any other, reading the same replicated state every agent reads, which is why it needs no instrumentation, no agent, no sidecar and no export: there is nothing to wire up because the coordination state is the telemetry.

See the fleet room for the operator's tour, and run examples/example-09-fleet-console to drive a live fleet: pause a worker mid-run and watch its queue hold.

Where it livesConsoleViewCustomizer, ConsolePanel and ConsoleCommand are @Beans collected by the console autoconfiguration, not agents — they need no @AgentSpec and get no card. They are plain beans on a peer where the console is off, so the same jar runs with and without it. The DirectiveGate is the opposite: it belongs inside the worker, because pausing is something a work loop does, not something done to it.
Keep in mindA directive gate only honors the console peer the deployment configured, narrowed further by the directive-issuer grant, and rejects future-stamped directives — so pausing a fleet is an authorized operation, not a message anyone can write into the space. The gate is a subscription: close it, or the lease keeps being renewed after the worker is gone. And @SpaceTake workers have no gate hook, so a fleet that must be pausable either writes its own loop as above or checks the gate inside the take method and returns early.
aspace://guide/gateways · chapter 8

Gateways and other languages

A2A. The gateway (agentspaces-a2a) serves every discovered AgentCard on the standard A2A discovery surface and speaks the task protocol as JSON-RPC over HTTP, with SSE streaming and push notifications: message/send becomes a leased task entry, tasks/get reads state from space semantics, message/stream and tasks/resubscribe stream over Server-Sent Events, and a lapsed lease honestly reports the task as submitted again. One gateway peer opens the door for the whole fleet, so individual agents never expose an inbound endpoint. It binds to loopback by default; discovery GETs stay open, the task surface takes a constant-time static bearer token or an identity-provider JWT carrying aspace:a2a:client, and push-notification webhooks validate against a private-address denial with an explicit allowlist override. Task results carry the authenticated issuer alongside whatever the worker declares about itself.

// One gateway peer fronts the whole fleet: it serves every discovered
// AgentCard as an A2A agent card, and binds the task protocol to a space.
try (A2aGateway gateway = new A2aGateway("research-fleet", "AgentSpaces demo",
                () -> discovery.find(AgentCard.class, card -> true))
        .taskBinding(new A2aTaskBinding(tasks))) {
    int port = gateway.start(8080);
    // POST /agents/{name}  {"method":"message/send", ...}  -> a leased task entry
    // POST /agents/{name}  {"method":"tasks/get",   ...}  -> state read from the space
}
Where it livesOn one peer, wired at startup and closed at shutdown — a @Bean with a destroy method, or a try-with-resources in a plain-Java peer. Individual agents know nothing about it and need no change to become reachable: the gateway serves whatever cards discovery is holding, so an agent joins the A2A surface by joining the fleet.

The connector SDK. agentspaces-connect-core is how a data source becomes a peer. An AssetProvider describes what it holds and answers queries for it; ConnectorRuntime publishes leased AssetCards and serves aspace:cap/data-query through the space, with DataQueryClient giving pull-once semantics keyed on a recomputed canonical query hash. MaterializingAssetProvider goes the other way, pushing source changes into a space as leased entries, and DataSpaces wires the block exchange so bulk results travel content-addressed. The SPI and the protocol are Apache 2; the module is marked EXPERIMENTAL.

// The provider peer publishes an AssetCard and answers queries for it.
try (ConnectorRuntime connector = new ConnectorRuntime(dataSpace, discovery, identity,
        "pg-connector", ordersTable, groupId, InstantSource.system(),
        Duration.ofMinutes(1))) {
    connector.start();
}

// Any analyst peer finds it by meaning, then fetches -- pull-once, so the
// second identical query anywhere in the fleet is served from the space.
DataQueryClient client = new DataQueryClient(dataSpace, "analyst-a");
DataQueryClient.Fetched page = client.fetch(card.asset(), Map.of("region", "eu"),
        Duration.ofSeconds(10)).orElseThrow();
page.result().rows();     // page.fromCache() is true when nobody re-ran the source

Python and TypeScript. The wire protocol is CBOR plus Ed25519 (wire version 2, enforced) with nothing Java-specific, and the project ships wire-compatible Python and TypeScript peers proven byte-identical against shared golden vectors that cover every wire structure, refusing the same hostile inputs as the Java implementation. Both join by seed or by GroupId with the founding document verified, write signed entries, take under the claim lattice, and read through anti-entropy. Both also speak the agent-identity model of chapter 6: they verify agent certificates, apply the two-key rule to records and take-claim proofs, can sign as a subordinate agent themselves (Identity.subordinate), carry agentPublicKey on the cards they build, and tally QUORUM ballots at either granularity (agentspaces.vote, vote.ts); their Peer takes an injectable clock so time-bound material is testable against a fixed instant, as Java's TestClock is. run-trilingual-demo.sh races workers in all three languages on one space.

Clojure. agentspaces-clj gives idiomatic bindings outside the Maven reactor, consuming the published artifacts by coordinates: maps in and out over record entry types, a keyword :where DSL, and fleet/start from a config map that mirrors the Spring starter's properties key for key. For the exact formats and algorithms per layer, continue to TECH-SPEC.md.

Where to go nextThe fourteen numbered examples add one layer each (examples/, from hello-space through the one-file quickstart to the exactly-once desk, the signed agents of example-13, and the gossip-learning fleet of example-14, each with a test that drives it end to end); the four flagship fleets under flagships/ are complete applications, standalone and movable to their own repositories; and agentspaces-partybus is the largest, thirty-four agents using every coordination type at once, with its own guide. The fleet room is the tour; this guide is the reference; SPEC.md and TECH-SPEC.md are the deep end.