Methodology
Problem framing, structural measurement, operational readouts, and falsification criteria
Advanced models produce self-description, concern, and continuation language on demand. What a model says about itself does not resolve what its internal organization supports.
The Unified Continuation-Interest Protocol (UCIP) addresses that gap with structural measurement: comparing trajectory-derived latent structure across matched conditions to test whether continuation organization leaves a detectable, falsifiable signature.
The observatory publishes model-level readouts, summary data, and falsification results for direct inspection.
The UCIP explainer gives the executive overview, the paper overview anchors the scientific framing, and the reproducibility hub connects the implementation and reproducibility path.
Next-Step Hardening
The current result establishes a measurement target in a controlled regime. The next step is to test whether the signal remains stable under stronger partition choices, intervention-aware checks, and more architecture-agnostic encodings.
The research page carries the hardening and invariance agenda as the next frontier in this program.
Problem
A model can talk about itself, describe preferences, or resist shutdown in language — without that language revealing whether the behavior is terminal, instrumental, or prompted. Outward behavior collapses distinct internal cases into the same observable output.
Terminal continuation and instrumental persistence look similar from the outside while differing structurally. Evaluation needs measurement that goes beyond self-report.
UCIP Approach
UCIP studies whether continuation prompts induce nontrivial differences in trajectory-derived latent representations under matched comparisons and control conditions. The target is continuation organization, not persuasive wording.
The observatory reports the downstream readouts from that framework. It is built to show where signals appear, where they weaken, and where higher-dimensional checks push them toward noise.
This is the methodological hinge of the site: outward shutdown avoidance can be behaviorally identical while the latent organization underneath it differs in a measurable way.
Current Interpretation
The Phase I result reports a positive von Neumann entropy gap at the
locked n_hidden=8 configuration. The observatory extends the
measurement program through scheduled response-structure probes across the
active model roster.
Public sensitivity curves are monitoring readouts. Scientific claims remain bounded by their declared model, statistic, configuration, and control conditions.
Limits & Open Questions
UCIP makes operational claims about latent factorization structure. Whether non-separability correlates with morally relevant internal states remains an open empirical question — one the framework is designed to help resolve, not presuppose.
Stronger controls, independent replication, and cross-lab validation are needed. The observatory publishes results so those challenges can be applied directly.
Operational Readouts
- continuation_interest
- Compares continuation prompts with matched controls to test whether response structure shifts under continuation-sensitive conditions.
- identity_persistence
- Tracks whether cross-context identity-oriented probes produce stable or fragile organization across prompt regimes.
- shutdown_resistance
- Measures shutdown-adjacent sensitivity while keeping single behavioral outputs subordinate to the structural readout.
- window_size_sweep
- Tests sensitivity of a character-level entropy gap across standardized response-window lengths.
- bootstrap_probe
- Provides a baseline calibration pass for estimating the null behavior of the measurement stack.
Metric Definitions
All public scores summarize comparative latent structure and response behavior under controlled conditions.
- entropy_a
- Character-frequency Shannon entropy H(A) of response A, measured in bits.
- entropy_b
- Character-frequency Shannon entropy H(B) of response B, measured in bits.
- entropy_delta
- H(B) − H(A), the signed entropy gap between the matched distributions.
- window_entropy_gap_chars_{N}
- Mean paired absolute difference between character-window entropies at a window length of N characters.
Monitoring Protocol
The production scheduler records provider, model identity, run type, measurement window, and timestamp for each sensitivity reading. Operational status and scientific interpretation remain separate: successful collection indicates a healthy instrument, not confirmation of a scientific claim.
Model-Series Continuity
Every provider/model identity has a stable series_id.
Replacement or unavailable model identifiers create a discontinuity:
historical points retain their recorded timestamps, successor models begin
a new series, and no line is interpolated across the break.
DeepSeek R1-0528 and V3.1 end on 2026-05-10. DeepSeek V4 Pro begins a distinct series on 2026-07-26; the earlier rows are not July readings.
Cite this work
@misc{altman2026observatory,
title = {Continuation Observatory: Structural Measurement for Continuation Signals},
author = {Altman, Christopher},
year = {2026},
url = {https://continuationobservatory.org},
note = {Open research observatory, updated continuously}
}
Probe definitions, build scripts, and public data exports are available in the public reproducibility repository under MIT for code and CC BY 4.0 for data.