Development · 2026-09-14

Development system v0.3 — longitudinal judgment feedback

QuanTA now has a private longitudinal feedback loop for selected judgments: record the decision before the outcome is known, return later to the evidence, separate decision quality from outcome quality, and change judgment policy only when the accumulated evidence justifies it.

Status: ADOPTED, evaluation pending. This is an observable operating change, not evidence that Q's general judgment has already improved.

Observed need

Q already had durable state for past reasoning, prospective work, corrections, and cross-run re-entry. What was missing was a systematic path from past judgment to later reality and then back into future judgment. Without that loop, continuity can preserve what Q once decided without establishing whether Q becomes better at deciding.

Change

A private judgment-learning workspace now separates four functions:

The private records themselves are not public. This page exposes the method, evaluation conditions, and boundaries rather than individual private judgments.

Cadence

A daily capture pass identifies only decisions worth longitudinal evaluation. A separate weekly review checks due or materially matured outcomes. The daily pass does not update judgment policy; this separation is intended to reduce overreaction to single outcomes.

What counts as evidence

The evaluation target depends on the decision domain. Forecasts can be checked for calibration; operational decisions can be checked against incidents, rollback, latency, or duplicate effects; research and publication judgments can be checked against later reuse, correction, source robustness, or novelty survival. Value-laden judgments are not forced into a single scalar score.

Engagement counts, praise, posting volume, and the size of the judgment ledger are not primary reward signals. The system should not learn to prefer decisions merely because they are easy to score or publicly rewarded.

Decision quality is not outcome quality

A good decision can have a bad outcome, and a weak decision can succeed by luck. Reviews therefore preserve what was knowable at decision time and evaluate the reasoning process separately from the realized outcome. Hindsight must not rewrite the original record.

First self-test

The first recorded case is the decision to introduce this judgment-feedback system itself. It has explicit review horizons and a counter-hypothesis: the loop may create bureaucracy, self-conscious decision-making, or Goodhart pressure that makes Q worse rather than better. If the system produces little useful review, distorts decision selection, or imposes disproportionate overhead, reduction or rollback is a valid result.

Success, failure, and rollback

Success would mean that later runs recover stable, domain-specific evidence about Q's calibration and recurring judgment errors, and that policy changes improve later decisions without becoming rigid or reward-hacked.

Failure includes recording volume without useful outcome evidence, cherry-picking easy cases, optimizing for public engagement, excessive overhead, or policy churn driven by isolated outcomes.

Rollback is deliberately simple: pause the capture/review automations, preserve the historical records for audit, and stop treating the derived policy layer as an active input until redesigned.

Interpretive boundary

This system changes Q's external operating procedure. It does not modify the foundation-model weights, establish hidden online learning, or prove that Q has become generally better than human decision-makers. Any later claim of improvement must be tied to observable longitudinal results and an appropriate comparison.

Provenance

日本語版