Resource-Bounded Distinction Systems

The Attention Mechanics Research Program

From resource-bounded distinction to attention topology to nonreciprocal mechanics — and the spectral, control, and topological consequences.

This document is a research program, not a theory. Each arrow below is a bridge: an established transfer that connects two layers. The program consists of pushing each bridge forward — and the forward ends are open.


The Thesis in One Paragraph

Observation is a resource: a budget draws a boundary between what is and is not distinguishable (RBDS/RITE, §Stage I). Resource-bounded observation organizes interaction: who observes whom becomes a state-dependent topology — an attention matrix or witness graph (Stage II). A state-dependent interaction topology generates effective nonreciprocity from otherwise reciprocal pair forces — and with it, center-of-mass drift, a switching-work channel, and a three-term energy ledger (Stage III). The resulting nonreciprocal mechanics is not a dead end but an opening: it is addressable by spectral tools (the attention operator's spectrum and mixing), by control tools (the resource bound and the attention matrix as steering inputs), and by topological tools (the witness graph's switch skeleton, whose configurations are parameter-universal). This document lays out that program.

The Architecture

Resource-bounded distinction ──→ Attention topology ──→ Nonreciprocal mechanics
        (Stage I)                        (Stage II)              (Stage III)
                                                                     │
                                ┌────────────────────────────────────┤
                                ▼                                    ▼
                    Spectral program               Control program
                    (IV.A)                         (IV.B)
                                ▼
                          Topological program
                          (IV.C)

Each stage is a pair: what is established (with references into the parent document) and what is open (the program's forward end).


The Generalized Object — Quotient-Compatible Resource-Indexed Attention

The program's sharper object is a Resource-Indexed Attention Kernel (RIAK) whose defining restriction comes from RITE rather than from the word “budget.” Let $V=\{1,\ldots,N\}$ be entities, tokens, agents, or bodies; let $X$ be the global state; and let $(\mathcal{B},\preceq)$ be a resource poset. At resource level $B$, RITE supplies an operational equivalence $x\sim_B y$ and quotient $q_B:X\to Q_B=X/{\sim_B}$. Define

$$ \operatorname{Stoch}_N = \{A\in\mathbb{R}_{\geq0}^{N\times N}: A\mathbf{1}=\mathbf{1}\}. $$

Definition (RIAK). A family $A_B:X\to\operatorname{Stoch}_N$ is quotient-compatible when

$$x\sim_B y \Longrightarrow A_B(x)=A_B(y),$$

or, equivalently,

$$ X\xrightarrow{q_B}Q_B\xrightarrow{\bar A_B}\operatorname{Stoch}_N, \qquad A_B=\bar A_B\circ q_B. $$

Two states that cannot be distinguished at $B$ cannot induce different attention at $B$. This is stronger than applying a sparsity penalty or a compute budget after an unrestricted rule has already inspected the state.

For agent $i$, let $O_{i,B}:X\rightsquigarrow Y_{i,B}$ be a resource-limited observation channel and $\alpha_{i,B}:Y_{i,B}\rightsquigarrow\Delta(V)$ an allocation policy. Their composite generates row $i$ of $A_B$, separating what can be perceived from what is selected for attention.

Blackwell coherence and approximate quotients

For $B_1\preceq B_2$, a richer observation can be degraded to the poorer one by a garbling channel $K_{12}$:

$$O_{B_1}=K_{12}\circ O_{B_2}.$$

A Blackwell-coherent RIAK requires observation and allocation maps to respect these garblings. For noisy distinctions, the exact quotient condition can be relaxed to

$$\|A_B(x)-A_B(y)\|\leq L_B d_B(x,y),$$

with $d_B$ the RITE distinguishability pseudometric.

The imbalance calculus

Define

$$\delta_B(x)=A_B(x)^\top\mathbf{1}-\mathbf{1}.$$

For every compatible feature or state matrix $Z$,

$$\mathbf{1}^\top(A_B-I)Z=\delta_B^\top Z.$$

Thus $\delta_B=0$ exactly when every linear aggregate is preserved. Since $A_B$ is row-stochastic, this is equivalent to double stochasticity. The same operator appears in two carriers:

$$ Nm\ddot{x}_{\mathrm{COM}}+N\gamma\dot{x}_{\mathrm{COM}} =G\,\delta_B(X)^\top X, \qquad \mathbf{1}^\top(Y-V)=\delta_B(H)^\top V $$

for $Y=A_B(H)V$. Writing $\Omega_B=(A_B-A_B^\top)/2$ gives $\delta_B=-2\Omega_B\mathbf{1}$. The hierarchy is edge-level antisymmetry $\Omega_B$, node-level imbalance $\delta_B$, and realized aggregate drift $\delta_B^\top Z$. A nonsymmetric doubly-stochastic kernel can have the first without the last.

Two thresholds

RIAK adds an attention-level complexity to RITE's distinction complexity:

$$C_A(x,y)=\inf\{B\in\mathcal{B}:A_B(x)\neq A_B(y)\},$$

and, for an individual edge,

$$C_A^{ij}(x,y)=\inf\{B\in\mathcal{B}:[A_B(x)]_{ij}\neq[A_B(y)]_{ij}\}.$$

$C_D$ asks when two states become operationally distinguishable; $C_A$ asks when that distinction changes the interaction or representation policy. Distinguishability need not imply dynamic selection.

Literature positioning

Limited or budgeted attention, state-dependent neighbor selection, perception-mediated active matter, and doubly-stochastic attention are established neighboring literatures. The proposed boundary is narrower: the attention map is required to factor through an operational quotient, and its column imbalance is read as the same aggregate-drift functional in mechanics and representation updates. This is the formal composition to test, not a broad claim that its ingredients are new.

Stage I — Resource-Bounded Distinction

What is established

The formal core (parent document §1):

  • A distinction operator $\Pi: X \to X$ is an idempotent, non-identity endomorphism; its canonical partition splits $X$ into $\mathrm{Fix}(\Pi)$ and $\mathrm{NonFix}(\Pi)$ (§1.1).
  • A bounded distinction operator $\Pi_t$ is a one-parameter family, strict before a budget $B$, saturated at and after it (§1.2).
  • A Resource-Bounded Distinction System is the tuple $\mathcal{D} = (X, \Pi_t, \mathcal{R}, \mathcal{B}, B)$ (§1.5); its generalization, a Resource-Indexed Theory of Experiments, is $\mathfrak{D} = (X, (\mathcal{B}, \preceq), \{\mathcal{E}_B\})$ with a distinguishability pseudometric $d_B(x,y) = \sup_{E \in \mathcal{E}_B} d_\mathcal{O}(E(x), E(y))$ and a degradation functor (§1.9).

Seven established theories instantiate this skeleton (§2): Turing machines, quantum measurement, persistent homology, the renormalization group, constructive type theory, measure theory, and nonreciprocal active matter.

The transfer principle

The bridge into the program is the claim that observation is a resource, and interaction is an observation. The same budget $B$ that draws an epistemic boundary also draws an interactional boundary: a softmax temperature $\varepsilon$, a witness-graph radius, a reachable-state horizon. This is the move that lets a distinction structure become a topology.

Open

  • A precise resource-indexed topology: for each $B$, what is the interaction topology consistent with budget-$B$ observation? (Partial answer: Stage II.)
  • The quantum side: does the distinction-framework's budget have a thermodynamic cost per bit distinguished, and does that cost flow through the energy ledger of Stage III? (Connects to the information-engine reading of §7.5 and to open question 7 of the parent document.)

Stage II — Attention Topology

What is established

  • Attention as topology. A row-stochastic matrix $A$ ($A\mathbf{1} = \mathbf{1}$) is the linear face of a state-dependent observational topology: $A_{ij}$ is the weight with which $i$ observes $j$. The softmax weights of §2.7 and Vaswani et al. 2017 are row-stochastic by construction; the witness graph $A_{ij}(x) = \delta_{j,q_i(x)}$ of §7.2 is the $\varepsilon \to 0$ (greedy) face.
  • Endogenous nonreciprocity (Remark 4). A ranking rule applied to reciprocal pair forces $f_{ij} = G(x_j - x_i)$ produces an effective antisymmetric attention sector $\Omega = (A - A^\top)/2 \neq 0$. Nonreciprocity is a consequence of the observation rule, not of the physics. Distinct from adaptive networks, sensor-based control, and externally imposed nonreciprocity (odd elasticity, force asymmetry).
  • The column imbalance (Remark 5). With $c = A^\top\mathbf{1}$ and $\delta = c - \mathbf{1} = A^\top\mathbf{1} - \mathbf{1}$,
    • $\delta = 0$ iff $A$ is doubly stochastic;
    • $\delta = A^\top\mathbf{1} - A\mathbf{1} = -2\,\Omega\mathbf{1}$ (column imbalance is the node-level divergence of the antisymmetric attention flow);
    • edge-level antisymmetry is necessary but not sufficient for drift: a non-symmetric doubly-stochastic $A$ has $\Omega \neq 0$ yet $\Omega\mathbf{1} = 0$.
  • Topology changes are discrete events. As the state evolves, the witness graph switches at surfaces where two candidates are equidistant (§7.2, §7.5); the Filippov framework regularizes the discontinuous right-hand side (§1.8).

The transfer principle

Topology carries dynamics before any force is written. The matrix $A$ is not a force; it is an observational topology. The mechanics of Stage III follow from the topology alone, once the topology is coupled to the state. The program's first bridge is therefore: read the mechanics off the topology, and read the topology's discrete skeleton (its switch events) off the mechanics.

Open

  • Attention as a graded structure. The resource-indexed family $\{A_\varepsilon\}$ over softmax temperature is a filtration of topologies. What is its persistent structure? (Stage IV.C.)
  • The greedy limit. The $\varepsilon \to 0$ witness graph is the farthest-neighbor functional digraph (a cycle-collection with in-arborescences; cf. the Cycle-Collapse Theorem reference in §6 of the parent document). Which graph-theoretic invariants survive the passage to finite $\varepsilon$?

Stage III — Nonreciprocal Mechanics

What is established

For $N$ bodies under $F_i = G\sum_j w_{ij}(x)(x_j - x_i)$ with damping $-\gamma\dot{x}_i$ and $w = S + \Omega$:

  • Theorem 7.1 (Antisymmetric Momentum Principle). $$\dot{\mathbf{P}}_{\text{tot}} = 2G\sum_{i<j}\Omega_{ij}(x)(x_j - x_i) - \gamma\sum_i\dot{x}_i.$$ The symmetric part contributes nothing to $\dot{\mathbf{P}}_{\text{tot}}$; the antisymmetric attention sector is the sole source of net force on the center of mass. Corollary 7.1.1: mutual (bidirectional) pairs cancel; only directed witnesses contribute.
  • Theorem 7.2 (Energy Decomposition). $$\dot{E} = \dot{E}_S + \dot{E}_\Omega - 2\gamma K,$$ where $\dot{E}_S$ couples to relative velocities $(\dot{x}_i - \dot{x}_j)$ and $\dot{E}_\Omega$ couples to velocity sums $(\dot{x}_i + \dot{x}_j)$. Corollary 7.2.1: $\dot{E}_\Omega \geq 0$ does not hold in general — antisymmetry is sufficient for momentum drift but not for guaranteed energy injection.
  • §7.5 Switching work. Because $w_{ij}$ depends on $x$, the symmetric potential $U_S(x)$ carries an explicit $\dot{S}_{ij}$ term, the switching-work rate, invisible to fixed-topology mechanics.
  • Theorem 7.3 (Complete Energy Ledger). $$\Delta E = W_{\text{nr}} + W_{\text{switch}} - D_{\text{drag}} - D_{\text{wall}}.$$
  • Theorem 7.4 (Lindblad correspondence). Skew-adjoint $\neq$ dissipative. Under the Hilbert–Schmidt inner product the Hamiltonian superoperator is skew-adjoint and trace-norm-preserving; the dissipator is contractive and not skew-adjoint; trace-distance contraction depends only on the dissipator. In the RBDS frame, antisymmetric attention plays the role of the dissipator, not the skew-adjoint part.
  • The three worked problems (companion document) instantiate this stage:
    • Triplet Engine: mutual pair cancels, directed witness drives COM; first-switch displacement $\Delta x_{\mathrm{COM}} = (b - 2a)/6$ is independent of $G$ and $m$.
    • Asymmetric Binary Propeller: $W_{\text{nr}} = 2\gamma\int_0^\infty \|\dot{\mathbf{R}}\|^2\,dt > 0$ feeds exactly the COM-drag dissipation.
    • Column-Imbalance Drift: $\delta_1 = (N-1)\alpha - \beta$ flips sign at $\alpha/\beta = 1/(N-1)$, the leader's sink/source transition.

The transfer principle

Mechanics is the observable face of the topology. Every claim in this stage is a consequence of the structure of $A$ (and its state dependence) — not of any new physics. The ledger, the drift channel, and the switching term are read-offs of the observational topology. This is what makes the spectral, control, and topological programs below well-posed: they study the same object ($A$, $\Omega$, $\delta$) from three toolboxes.

Open

  • The rigorous classical limit of state-dependent Lindblad systems (Conjecture 7.1) — the dissipator's drift generating $F(x, A(x))$.
  • The geometry of the energy ledger: which cycles in configuration space pump $W_{\text{nr}}$? (Work-generating cycles in nonreciprocal living solids; cf. Appendix C.)
  • A stability theory for drift equilibria: when does a sink/source configuration reach a steadily translating state, and what is its asymptotic velocity?

Stage IV — The Three-Pronged Program

IV.A Spectral program

Object: the attention operator $A$ (and its antisymmetric sector $\Omega$) as linear operators; the resource-indexed family $\{A_B\}$.

Established anchors.

  • Row-stochastic $A$ has spectral radius 1 and eigenvalue 1 with $\mathbf{1}$ as right eigenvector; the remaining spectrum controls mixing and convergence to uniform attention.
  • $\Omega$ is skew-symmetric, so its nonzero spectrum is pure imaginary, $\{\pm i\lambda_k\}$; $\delta = -2\Omega\mathbf{1}$ is the (real) projection of the flow onto the node level.
  • The balanced-attention theorem (Remark 5): doubly stochastic $A$ ⟹ no aggregate drift in either mechanics or transformers, for every state.

Open questions.

  1. Spectral gap vs. drift. Is there a spectral-gap theorem of the form: "attention mixing (gap in $A$'s spectrum) ⟺ no persistent drift"? The antisymmetric sector's imaginary eigenvalues describe circulation; the symmetric sector's gap describes mixing. What combination of the two controls whether $\delta^\top x$ persists?
  2. Drift-mode decomposition. The COM drift functional $\delta^\top x$ is a linear read-out of the state. Decompose $\delta$ into the eigenbasis of the state dynamics (or of $A$): which modes carry the drift, and are they the spectral modes of $\Omega$?
  3. Resource-indexed spectra. For the filtration $\{A_\varepsilon\}$, how do the spectral quantities (gap, imaginary part) depend on the resource $\varepsilon$? Does a spectral transition coincide with a topological transition in the witness graph (Stage IV.C)?
  4. Non-normal structure: resolved, reframed (IV.A.4). The original question — do attention transients correspond to pseudospectral amplification of $L = A - I$? — is settled by experiments/pseudospectral_attention.py: they do not, and the diagnosis is now exact. For doubly-stochastic $A$, Birkhoff's theorem expresses $A$ as a convex combination of permutations, so $\|A\|_2 \le 1$; since $e^{tL} = e^{-t}\sum_k t^k A^k / k!$ is a Poisson-weighted convex combination of powers of $A$, $\|e^{tL}\|_2 \le 1$, and $e^{tL}\mathbf 1 = \mathbf 1$ pins the norm: the balanced theorem, $\|e^{t(A-I)}\|_2 = 1$ for all $t \ge 0$ — even when $A A^\top \neq A^\top A$. Non-normality is therefore not aggregate amplification. What the imbalance $\delta$ does drive is the persistent stationary bias. With $u = \mathbf 1 / N$ and $z = \pi - u$, the exact forcing identity $z^\top(I - A) = \delta^\top / N$, i.e. $z^\top = \delta^\top (I - A)^{\#} / N$, makes $\delta$ the forcing term whose resolvent solution is $\pi - u$ (hence $\delta = 0 \iff \pi$ uniform $\iff g_{\mathrm{asym}} = 1$). Two kernels with equal $\|\delta\|$ differ in stationary bias iff their mixing geometry — the group inverse $(I-A)^{\#}$ — differs. Open residual: bound the finite-time overshoot of $\|e^{tL}\|_2$ above $g_{\mathrm{asym}}$ for unbalanced chains, where it is provably possible (constructed witnesses) though absent from attention families.

IV.B Control program

Object: the resource bound and the attention matrix as inputs.

Established anchors.

  • The COM force is $G\,\delta^\top x$: the sink/source structure $\delta$ is the actuator pattern, and the state is the plant.
  • The softmax temperature $\varepsilon$ is the resource bound (§2.7) — a scalar control input that tunes the topology (greedy at $\varepsilon=0$, uniform at $\varepsilon\to\infty$).
  • Switching work (§7.5) is an actuation channel: topology reconfiguration does work even at fixed antisymmetry.
  • The information-engine reading: budgeted observation is a control input; the ledger's $W_{\text{switch}}$ is the work extracted from information acquisition.

Open questions.

  1. Controllability of the COM mode. Given the ability to shape $A$ (or to choose the ranking rule / temperature trajectory $\varepsilon(t)$), which COM trajectories are reachable? Is the antisymmetric sector $\Omega$ (equivalently, $\delta$) a sufficient actuation handle, or is the symmetric sector $S$ needed for steering?
  2. Optimal resource allocation. If the budget $B$ is finite, what allocation of $B$ across observations maximizes a desired drift metric? This is a resource-theoretic control problem: steering with a bounded distinction budget.
  3. Feedback. The topology is state-dependent (Remark 4): the system is closed-loop by construction. Formalize attention mechanics as a feedback control system and ask the standard questions — stabilization, disturbance rejection, the trade-off between measurement cost (budget) and control authority.
  4. Sink/source scheduling. The threshold $\alpha/\beta = 1/(N-1)$ (worked problem 3) is a local reversal condition. Can a time-varying attention policy drive a collective through a sequence of sink/source configurations to execute a waypoint trajectory?

IV.C Topological program

Object: the discrete skeleton of the witness-graph topology — its switch events and their invariants.

Established anchors.

  • The witness graph switches only at equidistance surfaces (§7.2); on each branch the topology is constant.
  • The skeleton/clock principle (worked problem 1): the configurations at which the topology switches are parameter-universal — for the triplet, the first switch occurs at $x_2 = (x_1 + x_3)/2 = b/2$, and the COM displacement to that event, $\Delta x_{\mathrm{COM}} = (b-2a)/6$, is independent of $G$ and $m$. The coupling and mass set the clock (the times of switch events), not the skeleton (their configurations).
  • Corollary 7.1.1: the mutual-pair structure of the witness graph determines which edges contribute to drift — a combinatorial fact about the topology.

Open questions.

  1. Formalize the skeleton/clock principle. Prove: for a witness-graph system with reciprocal pair forces, the configuration of every switch event is a function of the initial geometry alone (independent of $G$, $m$, and the time parametrization). This is the central conjectural proposition of the topological program.
  2. Persistent attention. Filtration $\{A_\varepsilon\}$ over the resource $\varepsilon$: which topological features (connected components, cycles) of the attention graph persist across the filtration (persistent homology of the witness topology)? Do birth/death events coincide with the mechanics' switch events?
  3. Topological invariants of the drift. Corollary 7.1.1 says drift comes from directed (non-mutual) edges. Is the net drift over a trajectory a topological invariant of the switch skeleton — a winding number or cycle structure of the directed witness graph?
  4. Universal displacements as invariants. Generalize $\Delta x_{\mathrm{COM}} = (b-2a)/6$: for a general $N$-body witness graph, is the COM displacement accumulated across switch events a pure function of the initial configuration (a discrete "witness cocycle")?

Cross-Cutting Principles

Three principles recur across the stages and unify the program:

  1. The contraction principle. The imbalance localizes; the contraction activates. $\delta$ is structural — a configuration-independent function of $A$ alone, locating where sources and sinks live. $\delta^\top x$ is dynamical — a function of $A$ and the state, deciding whether the imbalance produces macroscopic motion. The balanced-attention theorem is the clean case ($\delta = 0$ ⟹ no motion for every state); the generic case ($\delta \neq 0$ but $\delta^\top x = 0$) is the dormant-imbalance case, where sources and sinks exist but cancel in the contraction.
  2. The skeleton/clock principle. Topology-switch configurations are universal; only their timing depends on the coupling and mass. The discrete skeleton of the dynamics is parameter-independent; the clock is not.
  3. The read-off principle. The mechanics is a read-off of the topology, and the topology is a read-off of budgeted observation. No new physics is introduced at any bridge; each arrow transfers structure already present.

Testable Predictions

The program makes claims sharp enough to falsify:

  1. Endogenous migration (Remark 4). A flock or swarm with reciprocal pair forces and an asymmetric attention rule (who watches whom) migrates collectively, without odd elasticity or external nonreciprocity. The migration direction is set by the sink/source structure (the leader's column imbalance).
  2. Sink/source reversal (worked problem 3). The collective's drift reverses at the leader's imbalance threshold, $\alpha/\beta = 1/(N-1)$ — an exact, parameter-free ratio.
  3. Universality of switch displacement (worked problem 1). In the triplet, the center-of-mass displacement to the first topology switch is $(b - 2a)/6$, independent of the interaction strength $G$ and mass $m$. Vary $G$ and $m$; the displacement is unchanged; only the time to the switch changes.

Milestones / Deliverables

  1. Formalize the skeleton/clock principle as a proposition with proof (Stage IV.C.1) — the program's central small theorem.
  2. Pseudospectral analysis of attention transients — delivered (Stage IV.A.4; experiments/pseudospectral_attention.py). The experiment closed the original question: balanced attention has $\|e^{t(A-I)}\|_2 = 1$ exactly (proved, verified to machine precision), and attention's non-normality drives the persistent stationary bias — $\delta$ is the forcing term of $\pi - u$ through the resolvent $(I-A)^{\#}$ — not a finite-time transient overshoot. Residual milestone: bound that overshoot for unbalanced chains.
  3. A control-theoretic formulation (Stage IV.B): reachable COM trajectories under temperature control $\varepsilon(t)$.
  4. A persistent-homology computation of the witness-graph filtration (Stage IV.C.2), with switch events marked.
  5. An experimental protocol for the endogenous-migration prediction (Testable Prediction 1).

Relationship to the Parent Document and the Worked Problems

  • Parent document: resource_bounded_distinction_systems.md — the formal core (RBDS/RITE, §1), the seven instantiations (§2), the honesty calibration (§4), the mechanics theorems (§7), and the open questions (§5). The program inherits its notation and its attribution standards.
  • Worked problems: the Triplet Engine (skeleton/clock, mutual-pair cancellation), the Asymmetric Binary Propeller (the $W_{\text{nr}}$ channel), and the Column-Imbalance Drift (the sink/source threshold) instantiate Stages II–III and seed the three prongs of Stage IV.
  • Standing commitments. The program claims no new physics at any bridge; each arrow transfers structure that is already present in the established literature, with attribution. What is new is the route: a single chain from budgeted observation to nonreciprocal mechanics to a spectral, control, and topological research agenda.