Symviosis
Symviosis Research

Load-Bearing Metrics

What to measure when building coordination infrastructure for systems change.

A research brief for systems architects, platform builders, and weavers — the connectors of projects, capital, and people. It covers which metrics a systems-change network should rally on, in what sequence, with which anti-gaming defenses, which immune-system metrics protect the network while it grows, and how to instrument emergence without destroying it.

Symviosis · Milan · 2026
Series: cognitive infrastructure · WP-03
§01 · Thesis

Measurement is architecture.

In a coordination network, whoever controls measurement controls the movement. Metrics decide what members see, what stewards optimize, where capital routes, and which behaviors become rational. Choosing rally metrics is a governance act with the same weight as choosing an ownership structure.

A systems-change network rallies on three metrics in a causal chain, instruments a wider set silently, defends every number against gaming, and keeps one shadow metric that can falsify the whole dashboard.

Rally, instrument, shadow

Rally metrics are public and legible to every member. Instrumented metrics feed the sensing layer silently. The shadow metric is checked and never rallied on — it exists to catch self-deception.

The chain over the portfolio

Three metrics that cause each other beat ten metrics that describe the network. Inputs are measurable and improvable; outputs can only be watched.

Adversarial by default

Every metric ships with its known attack surface and a defense. A metric published without its Goodhart analysis is an invitation.

The narrative is the story the metrics tell

"We measure whether we trust each other in ways that cost something, whether resources reach the edges fast, and whether dormant capacity becomes action." Legible to every member, and a public pre-commitment against the failure modes.

§02 · The builders

Three builder roles, one shared failure mode.

This brief addresses three roles that overlap in practice: network builders assembling coordination among people they know, platform builders shipping the infrastructure, and weavers routing projects, capital, and people across clusters. All three inherit the same structural risk at day one.

Role 01

Network builders

Convene hundreds to thousands of actors around a systems-change intent. Their asset is the trust graph; their risk is that the graph routes through them personally.

Role 02

Platform builders

Ship the coordination layer: telos intake, clustering, planning, commitment ledgers, telemetry. Their risk is building a central planner with better branding.

Role 03

Weavers

Connect projects to capital, people to projects, and clusters to clusters. Their risk is becoming invisible single points of failure that no dashboard tracks.

The founder-as-layer failure mode

At a few thousand personally-known members, the founder is the coordination layer. Algorithms sit on top of that trust graph — and if the founder disappears for three months, the network reverts to a contact list.

The first design job is transferring trust from the person to the protocol before scaling anything. The dependency never disappears on its own; it migrates to new hubs, including to the coordination tool itself. Measuring that migration is §16.

Scale facts to design with
  • 500 members ≈ 3× the Dunbar ceiling — small enough that trust density is still buildable through direct relationships, large enough that it requires deliberate construction
  • Past every Dunbar layer, status competition and free-riding are guaranteed, beyond merely possible
  • Clusters of 15–50 with named accountability, nested into larger units, work with primate biology
  • Coalescing visions means clustering by compatibility and negotiating interfaces between clusters — averaged visions motivate nobody
§03 · Coordination stack

The planner senses, proposes, and simulates. Humans commit.

A real-time AI that adjusts everyone's plans is a central planner with better branding, and it will get captured. Constraining its role is the architecture: the machine handles sensing, clustering, and simulation; humans hold every commitment decision; a contestability layer audits and constrains both.

FIG. 02Masterplanner reference architecture
Telos graph declared + revealed Signal layer commitments · capacity · context Sensemaking engine Cluster formation — compatibility, beyond consensus Plan compiler — proposals + simulations humans commit Commitment ledger — skin in the game Execution + outcome telemetry revealed data loops back Contestability layer audit · veto · fork Any cluster can see why the planner proposed what it proposed, veto it, and fork with their data if they lose trust. audits constrains forkability is the anti-capture guarantee: nobody fights to control a system they can cheaply exit
Two design items builders rarely price in.

A legal and data wrapper — ownership of the telos graph is a power question, well before it is a compliance question. And an opposition model — the moment the network coordinates real capital or political weight, incumbents attack it or build a controlled imitation. Communications and redundancy get designed for that day now.

§04 · Telos data

Declared telos is a hypothesis. Commitments are evidence.

Vision-coalescing that clusters on declared drivers inherits a systematic bias: people perform their purpose socially, especially high-agency people in systems-change contexts where purpose is status currency. Compatible-sounding clusters fracture on first contact with real resource allocation.

Revealed preference as first-class data

The correction: track what people commit time, capital, and reputation to, as first-class data alongside what they say. The telos graph carries both layers, and clustering weighs the revealed layer heavier as it accumulates.

  • Declared: intake answers, stated visions, chosen domains
  • Revealed: ledgered commitments, completed collaborations, staked reputation, capital moved
  • Divergence between the two is itself a signal worth watching
Consequence for every metric downstream

The declared/revealed split runs through this entire brief. Trust counts only when backed by costly signals. Activation counts only when it produces artifacts or ledgered commitments. Resource flow counts only when it reaches unique recipients. Every rally metric in §07 is defined on the revealed layer — the declared layer feeds hypotheses to test, never numbers to publish.

§05 · Selection problem

Rich frameworks measure everything. Networks rally on almost nothing.

Mature systems-change measurement frameworks run to dozens of macro categories and hundreds of sub-metrics across system perception, network architecture, agency, governance, resource flows, adaptation, culture, and biospheric domains. The richness is real, and it creates two problems for a builder at day one.

Problem 1 · Mixed ontological levels

Categories arrive at different taxonomic levels: system functions (information flow), outcomes (wellbeing), sectors (food), growth properties (horizontal diffusion), and methodologies (minimum viable evolving systems). Reasoning across them requires separation: system capacities on one axis, transformation domains on the other. The query pattern that results — "how does the regional food system perform on trust?" — is the coordination platform's actual data model.

Problem 2 · Rally capacity is scarce

Five hundred people can hold three numbers in their heads. Rallying on a 46-category dashboard rallies on nothing. The selection question is therefore brutal by design: which two or three metrics, given a starting scale of hundreds of people and a first fund of about a million, make the rest unfold — and which seductive candidates must be actively avoided.

The answer this brief defends: metrics forming a causal chain, one per load-bearing system layer, each scoped to a concrete transformation domain.

measurement engine = system layer → macro category → metric → observation → time state → node / cluster / system scale · queried per transformation domain
§06 · Criteria

What qualifies a metric to rally on.

A rally metric carries public weight: members judge the network by it, and the network judges itself. Six criteria separate rally metrics from everything that belongs in the silent sensing layer.

Upstream
It measures an input the network can act on, ahead of an output that can only be watched. Coordination is produced; its inputs are improved.
Falsifiable, fast
It can prove the architecture wrong cheaply. A metric that stays flat for 90 days must mean something decisive.
Legible
Every member understands it without a systems-thinking background. Composites and indexes fail this test structurally.
Costly to fake
It counts evidence over claims: artifacts, ledgered commitments, completed collaborations, capital moved.
Scoped
It attaches to a concrete transformation domain from day one. "Trust density in the regional food cluster" is measurable and improvable; "trust density in the movement" is a mood.
Defended
Its Goodhart attack surface is known, published, and countered before launch (§14).
§07 · Rally metrics

Three metrics forming a causal chain — never a portfolio.

Trust density, multicapital flow, and agency activation. Everything else in a systems-change framework — alignment, resilience, information flow, coordination capacity — sits downstream of these three or lags them. A network with high trust, fast resource routing, and visible shipped wins becomes magnetic: recruitment, alignment, and resilience stop being engineered and start accruing.

Rally 01

Agency & role activation

The rate at which dormant members become contributing members — measured weekly, at the network's edge. The binding constraint and the fastest-moving metric. Detailed in §08.

Rally 02

Trust, reputation & legitimacy

Cross-cluster trust ties backed by costly signals. The substrate every other metric runs on, and the one asset incumbents can neither buy, copy, nor regulate away. Detailed in §09.

Rally 03

Multicapital flow — velocity and reach

How fast resources move, and how far from the center they land. Converts a warm community into an operating network. Detailed in §10.

The causal logic

Trust density is load-bearing: information flows at the speed of trust, and coordination costs run inversely to it. Trust alone yields a warm community that does nothing — plenty of those exist. Resource mobilization makes pooled assets deployable in days. Shipped outcomes convert believers into committers and generate the revealed-preference data the telos graph needs.

Why inputs, why chain

The reverse direction fails: alignment and resilience resist direct optimization; only the conditions they emerge from can be built. Measure the inputs and the output arrives. Measure the output directly and diagnosis becomes impossible when it stalls.

§08 · Agency activation

The binding constraint is activated nodes.

At hundreds of members with a first fund on the table, the scarce resource is neither capital, information, nor ideas. A 500-person network typically carries 400+ spectators and dormant capacity worth ten times the fund — the retired water engineer sitting invisible in the community. Activation rate, dormant → contributing, is the single most falsifiable early metric: if it stays flat for 90 days, the architecture is wrong, and the network learns it cheaply.

Definition

What counts as activation

An activation event requires an artifact or a ledgered commitment with a deadline. Attendance, comments, and call participation carry zero weight — that is activation theater, the metric's first attack surface.

Segmentation

Edge activation is the number

The first activators are the people closest to the founder, which inflates early figures while the periphery stays dead. Segment by network distance from the core; the metric that matters is activation at the edge.

Uniqueness

New nodes over repeat nodes

The same forty people re-activating across projects imitates a healthy rate while laundering burnout. Count unique newly-activated nodes, and watch cognitive-load signals on repeat contributors.

§09 · Trust density

Bridging capital, backed by costly signals.

Bonding capital — trust inside clusters — forms on its own. Bridging capital across clusters is what makes 500 people a network instead of twenty cliques, and it is the Horizon-2 metric in transition terms: the bridge tissue between existing arrangements and emergent alternatives.

Measurement rule

Count only trust ties backed by costly signals: co-committed resources, completed collaborations, staked reputation. Trust claims are free; trust evidence carries cost. Cheap-reciprocity rings — mutual vouching, likes-as-trust — are the known Goodhart route and get excluded by construction.

Why it is unbuyable

Trust density is the one asset incumbents can neither purchase, copy, nor regulate away. It also moves on a months clock and resists direct construction — it gets harvested as the exhaust of completed activation events, which is why it rallies second, behind activation (§12).

§10 · Multicapital flow

Velocity and reach, measured together.

How fast financial, social, knowledge, and physical capital moves through the network — and how far from the center it lands. The reach component is structural: velocity alone can run high while everything flows to the founder's inner circle, which is clientelism with dashboards.

What the fund actually buys

With a first fund of about a million, the capital buys rails over outcomes: the commitment ledger and resource-routing infrastructure that make pooled resources deployable in days instead of months. A mobilization metric is fake when there is nothing to mobilize — the deployment pool has to exist for the number to mean anything.

Anti-gaming by topology

Circular flows — money moving fast between the same five nodes — score high on velocity while achieving nothing. Weight the metric by terminal impact and unique recipients, and audit flow topology rather than throughput alone.

§11 · Rejected candidates

The seductive metrics, and why each fails as a rally choice.

Four candidates dominate first drafts of every network dashboard. Each belongs somewhere — the sensing layer, the design desk, month twelve — and each fails as a day-one rally metric for a specific reason.

Rejected · Coordination capacity

The obvious pick is downstream

Coordination is what trust, activated agents, and flowing resources produce. Measured directly, it hides the diagnosis when it stalls. Measure the inputs; receive the output.

Rejected · Incentive architecture

A design variable, worn as a metric

Arguably the highest-leverage category in any framework — and at day one the builder sets the incentives rather than rallying 500 people to watch a number about them. It returns at month six as a drift detector between designed and actual payoffs.

Rejected · Narrative & memetic reach

The classic movement trap

Memetic spread is cheap, feels like momentum, and correlates with nothing material. Movements that rally on narrative metrics become content operations.

Rejected · Emergence potential score

Intellectually right, publicly wrong

A composite of 12–15 precursors, illegible to non-specialists, impossible to hold public accountability against. It belongs in the sensing layer (§18) — and the three rally metrics already cover three of its precursors: interaction frequency, excess resources, local agency.

Governance and resilience metrics wait, deliberately.

Governance measured early rewards performative process — participation theater before there is anything to govern. Resilience is a lagging property: at month zero there are no shocks to measure recovery from, so any resilience number is fiction. Both matter from month twelve onward.

§12 · Sequencing

The three metrics run on different clocks. Rally them in order.

Presented as parallel, the chain fails publicly. Activation moves in weeks — someone ships an artifact. Trust moves in months, harvested as the exhaust of completed collaborations. Resource velocity moves in quarters, because people route real capital only through trusted ties — the trust graph must bear load first. FIG. 01 draws the loop.

The public sequence
  1. Lead publicly with activation: fast, visible, falsifiable.
  2. Let trust accumulate as activation's exhaust; publish it once costly-signal ties exist to count.
  3. Open the resource rails once the trust graph can bear load; resource velocity then funds new activation, closing the loop.
The failure of simultaneity

Rallying on all three from day one buys ninety days of two flat metrics and a credibility problem. The slow metrics read as failure precisely when the architecture is working as designed. Sequencing is a communications decision with structural consequences: the network's belief in its own instrumentation is itself load-bearing.

§13 · Deficit diagnosis

Why trust is scarce and agency dormant — and why the split matters.

Running the causal diagnosis on the deficits themselves changes how each metric is expected to move. Among hundreds of capable people, dormant agency and mutual suspicion have two sources with opposite operational profiles.

Induced deficits — move fast

Extractive equilibria train passivity and suspicion actively: clientelist allocation, credential gatekeeping, and precarious labor keep people atomized because atomized people who distrust each other cannot pool procurement, bid collectively, or threaten anyone's rents. Dormant agency is the incumbent system's maintenance mechanism, beyond a bug.

The induced portion moves fast once rails exist. That is the first-90-days upside, and the reason edge activation can surprise positively.

Inherent deficits — get routed around

Kin-selection trust radii, Dunbar ceilings, and status competition move at zero speed. They get routed around with middleware: reputation systems and commitment ledgers functioning as trust prosthetics, extending cooperation beyond the radius biology grants.

This reframes what the coordination platform is: a prosthetic that lets 500 people behave as if they had 500 kin.

§14 · Goodhart defenses

Every metric ships with its attack surface and its counter.

A published metric is a target. Each rally metric carries a known gaming route; each defense is built into the metric's definition before launch, and published alongside it as a public pre-commitment.

MetricAttack surfaceDefense
Trust densityCheap reciprocity: mutual vouching rings, likes-as-trust.Count only ties backed by costly signals — co-committed resources, completed collaborations, staked reputation.
Resource velocityCircular flows between the same few nodes; high throughput, zero effect.Weight by terminal impact and unique recipients; audit flow topology, beyond throughput.
Agency activationActivation theater, founder-proximity inflation, burnout laundering.Require artifacts or ledgered commitments; segment by network distance; count unique new nodes.
Solution / conversion ratesShrinking ambition: commit only to trivial things and conversion looks stellar.Weight commitments and shipped wins by problem depth.
All metrics — the meta-attackWhoever controls measurement controls the movement; a planner that computes the metrics and allocates against them is the capture vector.The metrics pipeline gets the same contestability layer as the planner: auditable inputs, cluster-level right to challenge scores.
§15 · Shadow metric

Material access delta: the number that falsifies the dashboard.

The assumption that measurement drives movement fails in a specific, quiet way: all three rally metrics can rise for eighteen months while nobody's material life improves. High trust, flowing resources, activated agents — and rent unchanged, income unchanged. A beautiful dashboard for a social club.

Definition and custody

Material access delta: measurable change in what members can access or afford through the network — procurement savings, housing access, income routed, capacity unlocked. Checked always, rallied on never. It exists as a falsification instrument: flat at month twelve, it means the three metrics are measuring vibes, and the architecture restructures.

The second failure of metric-led rallying

Rallying on metrics is itself a telos filter: it selects for people who find dashboards motivating, skewing the network toward analysts and away from the builders and connectors telodiversity requires. The rally narrative is therefore the story the metrics tell — "a member's dormant expertise fixed the irrigation co-op's problem in three weeks" — with the numbers underneath, never in front.

§16 · Immune layer

Growth metrics measure the organism. Immune metrics defend it.

The deep bias in systems-change measurement frameworks: they are optimists' instruments. Dozens of categories measure how well the network grows; almost nothing measures what protects it while it grows. The asymmetry matters because failure runs faster than growth — trust built over eighteen months evaporates in one unhandled extraction scandal. Four immune metrics close the gap, and each exists because its absence kills a network in a specific way.

Immune 01

Exit & churn quality

Frameworks measure everything about presence and nothing about departure. Who leaves, at what network distance, and where they go next. Activated members leaving is damning evidence no dashboard of positive metrics can refute — and high-capacity people exit quietly, because voice costs more than exit. Unmeasured, healthy forking and silent bleed-out look identical. Movements die this way: dashboard green, kitchen empty.

Immune 02

Dependency & single-point-of-failure index

What fraction of trust paths, resource flows, and coordination routes pass through nodes whose removal fragments the network. Founder dependency migrates rather than resolving — to new hubs, to key weavers, and to the coordination platform itself, which is the dependency nobody volunteers to measure. This index tracks whether the architecture distributes or quietly re-centralizes.

Immune 03

Coercion & soft-power gradient

Trust metrics show whether people cooperate; nothing shows whether they can afford to refuse. When livelihoods, reputations, and access all route through one network, voluntary participation quietly becomes compulsory — the network reproduces the clientelism it was built to escape, with better aesthetics. Measure refusal cost directly, sampled through actual dissent events: can a member decline a request, dissent from a plan, or vote against a steward and keep material access?

Immune 04

Parasitism & extraction load

Accountability categories cover rule-breakers. This covers rule-followers optimizing take against contribution within the rules: drawing pooled resources, capturing introductions, spending collective reputation while contributing performatively. At scale, biology guarantees parasites; the only question is visibility of the load. Measure the drawn-to-contributed multicapital ratio per node, flagged at the tail — uncomfortable to measure, which is exactly why it is load-bearing.

§17 · Autoimmune risk

Immune metrics can attack the body they defend.

A network that measures parasitism, coercion gradients, and extraction load can tip into surveillance culture — everyone auditing everyone, trust corroding from the measurement itself rather than from the extraction it was built to catch. The defense is dosage and custody.

Dosage rules
  • Immune metrics run at aggregate and cluster level, never as public individual scores
  • They trigger human conversations, never automated consequences
  • Flags route to stewards with context, out of public view
  • Thresholds and methods are visible to all members, results are custody-limited
The line that must hold

The moment extraction load becomes a public leaderboard, the network has built a social credit system — and it deserves to be forked away from. Immune measurement earns its place by protecting cooperation; the instant it starts pricing individuals, it has switched sides.

§18 · Emergence

Emergence potential is a portfolio bet — never a KPI.

The precursor conditions are real: diversity × connectivity, unoccupied possibility space, recombination rate, latent complementarity, local agency, interaction frequency, boundary permeability, safe-to-fail capacity, excess resources, signal sensitivity, attractor instability, coordination latency, modularity, recursive amplification. Treating them as one more scored category misreads what they are.

The adversarial case against a precursor score

Most precursor conditions are also precursors of collapse. Attractor instability, unoccupied possibility space, excess resources, high boundary permeability — that list describes a system about to leap or a system about to disintegrate, and the precursors carry no information about which. A high emergence-potential score reports high variance; it stays silent about which tail arrives.

Goodhart on emergence is self-defeating

Optimizing diversity × connectivity directly produces engineered serendipity theater — curated collision events, cross-cluster mixers — which manufactures the measurable inputs while destroying the mechanism, because genuine recombination requires slack and unmanaged encounters. The moment a precursor becomes a goal, it stops being a precursor.

§19 · Event log + slack

Measure realized emergence retrospectively. Buy the substrate directly.

The allocator's approach replaces the dashboard's: track emergence as a portfolio of realized events with lineage, keep the precursors as silent diagnostics, and fund the one input that is boringly measurable.

01 · Event log

Realized emergence, with lineage

Log novel collaborations, solutions, and structures nobody planned — as events carrying lineage: which conditions, which nodes, which accidents preceded them. After 20–30 events the network holds its own empirical precursor signature, and only then do leading indicators earn trust.

02 · Instrument, target never

Precursors as diagnostics

The precursor set lives in the coordination platform's sensing layer — unpublished, target-free. It informs stewards where variance is building; it appears on no rally dashboard.

03 · Slack

The substrate metric

Percentage of network time, capital, and attention unallocated to committed plans. Measurable, resistant to Goodhart (an absence performs poorly on stage), and directly purchasable. Zero-slack networks can only execute; emergence lives in the unallocated remainder. The capital plan carves slack out explicitly as unrestricted, cluster-discretionary capacity.

Worked example · event record schema

event: cross-cluster water-sensing collaboration, unplanned. lineage: two clusters sharing a steward; slack capital drawn without proposal; a dormant member activated three weeks prior; boundary permeability with a research institute. precursors present: latent complementarity, local agency, excess resources. outcome: shipped prototype, two new costly-signal trust ties, one new watcher created.

§20 · Differential visibility

The trust graph is also a targeting map.

Visible topology tells whoever wants to capture the network exactly which bridge nodes to buy, flatter, or fund — the cheapest capture vector there is, and precisely how clientelist systems have absorbed civic movements for decades. Measurement infrastructure therefore ships with a visibility policy, designed with the same care as the metrics.

Aggregate metrics
Public. The rally numbers, network-level and per transformation domain, visible to every member and shareable outside.
Topology
Held at cluster level. Clusters see their own graph and their interfaces; the full map stays custody-limited.
Bridge-node identity
Protected by default. The highest-leverage connectors are the highest-value capture targets; their position in the graph is treated as sensitive infrastructure.
Immune metrics
Aggregate and cluster custody only, per §17 — human review, automated consequences never.
Metrics pipeline
Contestable end to end: auditable inputs, inspectable computation, cluster-level right to challenge any score (§14).
§21 · Operating model

Scale, capital, scoping, and the 90-day falsification test.

The measurement architecture compresses into an operating model a builder can run from day one, at a starting scale of roughly 500 members and a first fund of about one million.

Capital · ~40%

Coordination rails

Commitment ledger, resource-routing infrastructure, telos graph, metrics pipeline with its contestability layer.

Capital · ~35%

Rapid-deployment pool

Capital the clusters can actually draw on — the mobilization metric is fake when there is nothing to mobilize.

Capital · ~20%

The human layer

Cluster stewards and the in-person container. Trust density at this scale is built face to face; the platform carries it afterward.

Capital · ~5%

Slack carve-out

Unrestricted, cluster-discretionary capacity — the explicit purchase of emergence substrate (§19).

The launch checklist
  1. Scope each rally metric to one concrete transformation domain — a regional food system, a housing cluster, an energy commons.
  2. Publish the three metrics with their Goodhart defenses attached, as a public pre-commitment.
  3. Rally on activation first; instrument trust and resource flow silently until they carry signal (§12).
  4. Stand up the four immune metrics under §17 custody rules, and the shadow metric under founder-plus-steward custody.
  5. Open the emergence event log at day one; review the precursor signature at event 20.
The falsification contract

Edge activation flat at day 90: the architecture is wrong — restructure the intake, the cluster design, or the commitment mechanism. Material access delta flat at month twelve: the metrics are measuring vibes — restructure the model. Extraction flags rising with exit quality worsening: the immune system found something — act before the dashboard shows it.

A measurement architecture earns trust by naming, in advance, the numbers that would prove it wrong.

Closing thesis

Instrument the growth. Defend the organism. Leave room for the leap.

Three rally metrics in a causal chain — agency activation, costly-signal trust, multicapital flow — sequenced by clock speed and scoped to real domains. Four immune metrics watching exit, dependency, coercion, and extraction under strict custody. One shadow metric holding the whole dashboard falsifiable. Precursors instrumented in silence, emergence logged with lineage, and slack purchased deliberately. Measurement built this way transfers trust from founders to protocols — and gives a systems-change network the rarest property in the field: the ability to know, cheaply and early, when it is wrong.

SymviosisLoad-Bearing Metrics · Research brief
Designed for frontier systems intelligence.