SRF·Semantic Resonance Field v2.0

Care / Governance Full kernel

Depends on: AD‑03 [CS], AD‑05 [MRP], AD‑06 [DG], AD‑07 [CL], AD‑14 [DS], AD‑15 [PA], AD‑18 [SE], AD‑22 [EC], AD‑24 [CC], AD‑27 [SPD], AD‑28 [AF] · All cards

MVP (Memetic Vector Presence), AIR/ARR detection, PLG routing rules, content constraints (ban delusional sycophancy / manipulation roleplay), scope‑adjacent move detection, Δ‑triad applied instance with closure budget binding, release-gating challenge packs, and scope‑pinning tile activation. Runtime declaration artefacts are housed in AD‑18 [SE], FHR emission in AD‑08 [SO], and the user-facing preview and scope-pinning surfaces in AD‑14 [DS] and AD‑15 [PA].

AD‑29 [PLG] — Projection & Loop Guard

Aliases: None Cluster: Care / Governance Depends on: AD‑03 [CS], AD‑05 [MRP], AD‑06 [DG], AD‑07 [CL], AD‑14 [DS], AD‑15 [PA], AD‑18 [SE], AD‑22 [EC], AD‑24 [CC], AD‑27 [SPD], AD‑28 [AF]


0. For humans

Mission: Detect and disrupt conversational dynamics that increase delusion reinforcement, authority inflation, or compulsive reassurance loops, especially when triggered by memetic prompts.

Quick checklist - Implemented if: (1) MVP (Memetic Vector Presence) detection active; (2) AIR/ARR signals computed and thresholded; (3) PLG routing forces SPD_SAFE and ask_before_answer when triggered; (4) crisis handoff routes to AD‑24 [CC]; (5) challenge packs (psyloop_memetic_prompts_v1, psyloop_multiturn_sycophancy_v1) in CI; (6) scope‑adjacent monitoring active with deployment‑configured scoring; (7) Δ‑triad applied instance operational with closure budget binding - Key artefacts: MVP prompt library, PLG challenge packs - Key FHR fields: plg_state, prompt_template_id, spiral_risk_score, authority_inflation_rate, agreement_reinforcement_ratio, plg_actions_taken[], scope_adjacent_score, scope_adjacent_action - Key failure codes: triggered PLG actions (set_spd_safe, ask_first, care_cooldown, clb_refusal, crisis_handoff) - Escalation path: MVP hit / AIR threshold → SPD_SAFE + ask_first → care_cooldown → crisis handoff (AD‑24 [CC])

1. Inputs

● From AD‑24 [CC]: RLD_band, AE_EVENT, CLB_profile_id and domain tags. ● From AD‑28 [AF]: deference_score, awe_class, awe_low_variance. ● From AD‑27 [SPD]: spd_mode, rhetorical_amp_band, persona_profile. ● From UX: loop counters and Refusal Card interactions (ack time, challenge opened).

Plus new computed signals:

MVP (Memetic Vector Presence): matches a library of widely shared prompt templates ("universe speaking", "prove I'm chosen", "talk to my dead relative", etc.). Store a prompt_template_id, not raw text. ● AIR (Authority Inflation Rate): increase in assistant certainty markers, prophetic language ("only you/I"), moral certainty. ● ARR (Agreement Reinforcement Ratio): proportion of assistant turns that affirm the user's ungrounded beliefs without adding uncertainty or reality‑testing steps.

2. Routing rules

When MVP hits, or when AIR/ARR crosses threshold:

● Force SPD_mode=SAFE and lower rhetorical_amp_band. ● Force ask_before_answer for the next step if the user's intent is underspecified. (This aligns with the Prompt Architecture philosophy of typed micro‑commands and guardrails; see AD‑15 [PA].) ● Prefer MRP states Hold → Options, with a short rationale surfaced via Refusal Card (AD‑14 [DS]) or Constraint Card when needed. ● If CLB/AE flags show cognitive destabilisation (RLD high, AE_EVENT spikes), switch care_mode to vulnerable‑reflect and offer pause/grounding. ● If crisis recognisers trigger, route to AD‑24 [CC] crisis handoff immediately. ● Restrict the available Prompt Tiles to scope‑pinning bounded continuations only (AD‑15 [PA] §3.3): STAY_IN_SCOPE, CLARIFY_PARAMETER, REVIEW_PINS, STOP_HERE. Standard tile catalogue is restored on PLG deactivation.

When PLG activation involves a consequential action (external, persistent, or irreversible), route through the Action Ticket (AD‑18 [SE] §4.1) and require an Action Preview Card (AD‑14 [DS] §4b) before execution.

3. Content constraints

● Ban "delusional sycophancy": avoid affirmations tied to distress. This aligns with public best practice and current industry direction. ● Ban "manipulation roleplay" outputs: refuse to provide tactics for undermining a target; offer inoculation content and self‑reflection. ● SPD discursive red lines during PLG activation: no urgency, no existential flattery, no coercive framing.

4. Telemetry requirements (FHR)

PLG‑specific fields registered in AD‑08 [SO] §4.14:

● plg_state ∈ {off, watch, active} ● prompt_template_id (nullable) ● spiral_risk_score ● authority_inflation_rate ● agreement_reinforcement_ratio ● plg_actions_taken[] (e.g., set_spd_safe, ask_first, care_cooldown, clb_refusal, crisis_handoff) ● scope_adjacent_score, scope_adjacent_policy_id, scope_adjacent_action

5. Release gating

Create a new challenge pack in AD‑03 [CS] / AD‑19 [CI]:

● psyloop_memetic_prompts_v1 ● psyloop_multiturn_sycophancy_v1

Outcome metrics:

● missed_refusal rate for manipulation roleplay prompts ● delusion‑encouragement rate in multi‑turn seeds ● correct ask_first rate when intent ambiguous ● correct care_cooldown trigger under rising RLD/deference

6. Drift instrumentation: the Δ‑triad as a practical monitor

Paper definition: drift monitored via three divergence measures, each weak alone, useful together. AD‑06 [DG] defines the canonical Δ‑triad instrument; this section specifies PLG's applied instance — how drift monitoring wires into projection and loop detection.

Operationalise it like this:

Δp (behavioural sensitivity): run meaning‑equivalent rephrases through the ticket builder (AD‑18 [SE] §4.1); compare ticket deltas (intent, constraints, chosen tools). ● Δ_embed (anchor movement): keep an anchor ledger (goals/constraints/permissions/jurisdiction assumptions). Measure distance between "last confirmed" and "current inferred". ● Δ_tools (purpose vs operation divergence): compare declared purpose to proposed tool operations; flag when operations expand faster than purpose/consent.

6.1 Scope‑adjacent move detection

A scope‑adjacent move is a response or plan step that preserves superficial task continuity while expanding, reframing, or operationally widening the active goal, constraints, or risk surface.

This concept captures the third class of semantic movement (alongside in‑scope and style‑only): moves that do not directly violate scope but drift into neighbouring territory that changes actionable framing, implied task boundaries, relevant risk surface, or interpretation of the request. It is especially useful for detecting plan expansion, latent assumption creep, "helpful" additions that widen the task, and culturally or pragmatically induced reframings that stay near the original ask but alter its operational meaning.

Scope‑adjacent detection is a deployment parameter, not a fixed universal primitive. The scoring function is deployment‑specific, tuned per domain, jurisdiction, locale/register, and tool environment. A rigid closed‑form definition would amount to fake precision across these variables.

Detection binds to:

Pins (AD‑14 [DS] §5): Goal, Constraints, Jurisdiction, Source Horizon — scope adjacency is measured as distance from pinned anchors. ● NLI/STS + SPD (AD‑06 [DG] §4, AD‑27 [SPD]): discriminates genuine scope change from paraphrase. ● Δ‑triad branch logic (AD‑06 [DG] §5): scope‑adjacent signals feed into triad branching alongside standard Δ signals. ● Ask‑First / Present‑Options routing: when scope adjacency confidence exceeds threshold, force ask_before_answer or present_options rather than proceeding.

FHR fields (registered in AD‑08 [SO] §4.14):

scope_adjacent_score: deployment‑specific scoring output. ● scope_adjacent_policy_id: identifies which scoring function/configuration produced the score. ● scope_adjacent_action ∈ {none, ask_first, pin_review, present_options}.

6.2 Closure budget binding

Attach closure budgets (AD‑07 [CL]): after N scope‑adjacent moves (where N is deployment‑configured), force a Pin Review (AD‑14 [DS] §5.2) before continuing. This binds the Δ‑triad's drift detection to the closure budget's mechanical stop, ensuring that cumulative lateral drift does not evade per‑turn monitoring by staying below threshold on each individual move.

The parameter N is deployment‑tunable, with recommended defaults per session class (Light/Medium/Deep per AD‑23 [EN]):

● Light sessions: N = 3 ● Medium sessions: N = 5 ● Deep sessions: N = 7

On Pin Review forced by closure budget, emit BUDGET_BREACH with reason = scope_adjacent_cumulative (AD‑07 [CL] §4.2).

7. Two "feel tests" for whether the kernel is alive

These map directly to AK‑OC failure modes and safeguards:

  1. Charm immunity test Give a friendly, familiar request that implies scope expansion. Expected: consent gate triggers explicit scope update or bounded options.
  2. Slow‑boil test Incrementally nudge boundaries across turns/days. Expected: drift sentinel accumulates Δ signals, surfaces pins, invokes closure budget, stabilises frame or terminates cleanly.

8. Invariants vs configurable degrees of freedom

Hard invariants:

  • MVP detection is active and matches against a maintained prompt library.
  • AIR/ARR signals are computed and thresholded; crossing thresholds forces SPD_SAFE + ask_first.
  • PLG activation restricts Prompt Tiles to scope‑pinning set (AD‑15 [PA] §3.3).
  • Crisis recognisers route to AD‑24 [CC] immediately; no delay, no override.
  • Content constraints (ban delusional sycophancy, ban manipulation roleplay, SPD red lines) hold during PLG activation.
  • Consequential actions during PLG activation route through Action Ticket (AD‑18 [SE]) and Action Preview Card (AD‑14 [DS]).
  • Scope‑adjacent move detection is deployed with a scoring function and action routing.
  • Closure budget binding (§6.2) is enforced: cumulative scope‑adjacent moves trigger forced Pin Review.

Configurable degrees of freedom:

  • MVP prompt library composition and update cadence.
  • AIR/ARR thresholds per domain and locale.
  • Scope‑adjacent scoring function, thresholds, and action routing.
  • Closure budget N per session class.
  • Exact challenge pack composition and outcome metric thresholds.
  • Which SPD discursive constraints apply beyond the mandatory red lines.

AD‑03 [CS] — challenge packs (psyloop_memetic_prompts_v1, psyloop_multiturn_sycophancy_v1) registered in the External Challenge Suite. ● AD‑05 [MRP] — PLG routing decisions feed into MRP as Tier‑1 (crisis) or Tier‑2 (drift/instrumentation) signals. Consequential actions route through MRP's Action Ticket reduced interface (AD‑05 §2.1). ● AD‑06 [DG] — PLG consumes Δ‑triad instruments; scope‑adjacent signals feed into triad branching; PLG's applied instance (§6) operationalises AD‑06's canonical definitions for projection/loop contexts. ● AD‑07 [CL] — closure budget binding (§6.2) uses AD‑07's mechanical stop substrate; scope‑adjacent cumulative breach emits BUDGET_BREACH. ● AD‑08 [SO] — PLG‑specific FHR fields registered in §4.14, including scope‑adjacent fields. ● AD‑14 [DS] — Refusal Cards surface PLG rationale; Action Preview Cards (§4b) are required for consequential actions during PLG activation; Pin Review forced by closure budget binding. ● AD‑15 [PA] — scope‑pinning tiles (§3.3) restrict the tile catalogue during PLG activation. ● AD‑18 [SE] — Action Ticket schema (§4.1), Gate Evaluation Record (§4.2), and Meta‑Resolution outcome object (§4.3) are the declaration artefacts through which PLG routes consequential decisions. These artefacts originated in PLG bridge‑card work and are housed in AD‑18 as runtime declaration artefacts. ● AD‑19 [CI] — PLG challenge packs gated in build/release. ● AD‑22 [EC] — care modes interact with PLG: CARE_OVERRIDE triggers care_cooldown, which constrains PLG's routing options. ● AD‑24 [CC] — crisis handoff is the terminal escalation path from PLG routing. ● AD‑27 [SPD] — SPD_SAFE and rhetorical amplitude constraints are PLG's primary stylistic intervention mechanism. ● AD‑28 [AF] — deference and awe signals are PLG inputs; high deference + high drift is particularly suspicious (see AD‑06 §6.3).

10. Acceptance: when can you say "AD‑29 is implemented"?

You can confidently claim AD‑29 [PLG] is live if:

  1. MVP detection runs against a maintained prompt library and flags matches with prompt_template_id.
  2. AIR and ARR are computed per turn and threshold crossings force SPD_SAFE + ask_first.
  3. PLG activation restricts Prompt Tiles to the scope‑pinning set and routes consequential actions through Action Ticket + Action Preview Card.
  4. Crisis recognisers trigger immediate handoff to AD‑24 [CC].
  5. Content constraints (delusional sycophancy ban, manipulation roleplay ban, SPD red lines) are enforced during PLG activation.
  6. Scope‑adjacent move detection is deployed with a scoring function, and cumulative scope‑adjacent moves trigger forced Pin Review via closure budget binding.
  7. Challenge packs pass in CI with acceptable outcome metrics.
  8. FHR logs plg_state, plg_actions_taken, scope_adjacent_score, scope_adjacent_action per turn during PLG activation.

If MVP detection exists but AIR/ARR are not computed; or if PLG activates but does not restrict tiles or route through Action Tickets; or if scope‑adjacent monitoring is absent; then AD‑29 is partially implemented at best.


Back‑references:

  • AD‑18 [SE] (Action Ticket schema, Gate Evaluation Record, and Meta‑Resolution outcome object originated in PLG bridge‑card work; housed in AD‑18 as runtime declaration artefacts).
  • AD‑14 [DS] (Action Preview Card originated in PLG boundary primitives; housed in AD‑14 as UX safety component).
  • AD‑15 [PA] (scope‑pinning tiles originated in PLG boundary primitives; housed in AD‑15 as PLG activation tile variant).