DSL DIGITAL SIDEWALK LAB
SYS: ENT-EDU > ACTIVEFLD: CYBERNETICS > ONLAB: CONSOLE > LIVEPUB: IEEE 2025 > INDEXEDJRN: 33 > INDEXEDOSS: APACHE-2.0 > PENDINGSTS: OPERATIONAL
REGISTER · 2026-09-15

Journal — Method

What should we measure next?

An observation-design instrument, its one frozen failure, and how a failure became the calibration.

Work periodJune–August 2026
EvidenceValidated on synthetic agents; shelved pending real use
Published2026-09-15
AuthorDarby Bailey McDonough, Ph.D.

Given a body of evidence and several competing explanations, which measurement would most cheaply tell the explanations apart? That is the question the observation-design ledger is for, and in summer 2026 the Lab built an instrument to answer it. Its working name is Futures Sidewalk.

The instrument

The input is a set of evidence records and a set of competing hypotheses. The instrument runs a synthetic wind tunnel: a simulation of how each hypothesis would show up under each candidate measurement. From that it computes a minimum-cost cover, the cheapest set of measurements such that every pair of hypotheses is distinguished by at least one of them. The output is a ranked proposal for what to measure next, with its cost and the pairs it separates.

Everything it produces goes in the simulation and observation-design ledgers. Nothing it produces is evidence.

The frozen failure

The instrument had a set of acceptance tests. One of them, a count-based check, failed on a frozen run: ten flagged items against a ceiling of nine. The Lab's rule is that frozen failures are preserved, not re-run until they pass. So the failure was kept, and the question became whether the ceiling was justified.

It was not. The ceiling had been set by hand. The fix was a rule that every fixed ceiling must be calibrated from the empirical distribution of its own test under a null model, one in which nothing is happening. Two hundred null replicates were run. The count's null distribution had a mean near 4.7 and a 95th percentile of exactly nine. The hand-set ceiling had been numerically right and previously unjustified, and the original failure was adjudicated as a four-percent tail event. A fresh run at an untouched random seed passed, and that phase closed on August 11, 2026.

Independent reproduction

A second assistant reproduced the failing run exactly, including the failure, from the archived configuration, and corrected two of the Lab's own procedural errors along the way: a diagnostic run without archived manifests was reclassified as notes, and the initial null-replicate count was raised because it was too small for a 95th-percentile gate. Both corrections were accepted and are in the record.

Status

The instrument is validated on synthetic agents and is shelved. The Lab's rule for reopening it is that a real ambiguity in one of its own nodes, a case where two explanations of a real record cannot be told apart from existing data, has to call for it. It is not to be demonstrated on a manufactured example.