GROUP MEETING · 2026-07-24

XMM Proposal Agent

Turn a proposal idea into a scientific question that is
auditable, falsifiable, and able to survive rebuttal

41live ideas
5graveyard records
2developing ideas
1dedicated HEAVY lane

XMM Proposal Agent · internal research system · static snapshot 2026-07-23

01 · WHYThe problem is not prose generation; it is scientific selection.

Failure modes of the conventional proposal workflow

Topic first

Then search for evidence that supports it; archive duplication and published work are often discovered only at the end.

One agent self-validates

Proposing, retrieving evidence, writing, and judging are mixed in one context, inviting attachment.

Numbers added late

Count rate, background, pile-up, and systematics surface only late in the draft.

DESIGN PRINCIPLE

Treat every idea as an experimental object that can fail.

It must leave behind sources, evidence, rebuttals, unresolved gates, a feasibility chain, state transitions, and an owner for the next step.

idea = claim + evidence + attack + feasibility + state
02 · CORE CONCEPTA proposal is a terminal event; evolution is the steady state.

Move from “What can we write?” to “What is worth continuing?”

COMMON PATH
Topic
Draft
Late review
Patch

Narrative first; failure often arrives too late.

XMM PROPOSAL AGENT
Archive fact
Opportunity
Idea
Attack
Gate
Proposal

Evidence first; failure is retained as reusable knowledge.

“A killed idea remains knowledge.”
03 · ARCHITECTUREThree boundaries prevent chat logs from masquerading as a scientific process.

Separate role specifications, scientific memory, and execution traces

A

Stable role specifications

pipeline/*.md defines what each role can and cannot do; roles change rarely.

ROLE GRAPH
B

Accumulating scientific memory

Coverage matrix, TRUST_TABLE, idea ledger, decision log, and graveyard.

CANONICAL RECORD
C

Disposable but auditable traces

Per-cycle materials, critiques, rebuttals, calculations, and SHIFTS handoffs.

REPLAYABLE TRACE
04 · ROLES11 responsibilities + two execution lanes

Unbundle discovery, verification, imagination, attack, calculation, writing, and judgment

M0Material collectordiscovery ≠ trust
R0Surveyorarchive → map
R1Librarianquestion → facts
R2Proposeropportunity → idea
R3Analystidea → evidence
R4Red-teamidea → kill shots
R5Judgeevidence → state
R6Feasibilityclaim → numbers
R7Drafterrecord → draft
R8TAC-sim paneldraft → review
LTutorunknown → learning
Fresh-context requirement: the R2 proposer, R4 red-team, and R5 judge work in independent shifts; the same author does not judge their own idea.
05 · MODE 1Default: the continuous Evolution Cycle

Every cycle requires new external input; stop when there is none

  1. New inputR0 expands archive coverage, R1 ingests a new paper or data release, or the system consumes a human annotation.
  2. R2 · ProposeGenerate a baseline / fork from opportunity tags × science questions.
  3. R3 + R4Collect evidence; an independent red-team attacks every live idea.
  4. R6-liteOrder-of-magnitude sanity check for flux / extent / background.
  5. R5-PairwiseBounded Elo debate within the same track: compare rather than absolutely score.
  6. R5-PortfolioMin-gate; kill / park / promote; blocking item; INDEX.
  7. Handoff + logSHIFTS, CYCLE_LOG, and REVIEW_QUEUE become inputs to the next cycle.
06 · MODE 2Entered only when an AO is near and a proposal-ready idea exists.

Proposal Run is a constrained process with admission criteria

1

Freeze question

Create a versioned run from a proposal-ready idea.

2

R1 deep-dive

ObsID-level archive, duplication, and instrument facts.

3

R6 full

Model → rate → pile-up → background → exposure.

4

R7 draft

Write only within the scope already closed by evidence and feasibility.

5

R8 TAC-sim

Three heterogeneous reviewers; any poor dimension triggers revision / park.

Current state: 0 proposal-ready · the system therefore must not pretend to be a “proposal machine ready to write.”
07 · IDEA LEDGERA state is not copy; it is a traceable decision.

An idea’s lifecycle and the non-compensatory min-gate

seedAn opportunity exists, but a key discriminator is unfinished.
developingA falsifiable sub-claim remains alive.
proposal-readyTop Elo for ≥2 cycles + no FATAL + R6 pass.
draftedR7 + R8 pass.
DECISIVENESSgood / adequate / poor
TIERgood / adequate / poor
COST-EFFECTIVENESSgood / adequate / poor
UNIQUENESSgood / adequate / poor
08 · CASE STUDY AEven an idea that looks exciting must die early when it should.

I001: M82 time-domain superwind → killed by physics and systematics

CYCLE 2 · R2

Initial idea

Use roughly 20 years of XMM + Chandra archives to search for temporal changes in the diffuse M82 superwind, and connect them to the central engine.

On the surface: deep archive, multiple epochs, a writable story.

R4
ATTACK
KILLED · CYCLE 2

Three kill shots

  • Physics: sound-crossing time ~105 yr ≫ 20 yr baseline.
  • Decisiveness: cross-calibration systematics ~5–10% ≫ expected signal.
  • Instrument: ACIS contamination can mimic temporal variation.
Lesson: “Multi-epoch archive” is not sufficient for time-domain science; first ask about the physical timescale and the systematic-error floor.
09 · CASE STUDY BDuplicated work is not a minor issue; it is a termination condition.

I003: New XMM halo pointing for NGC 891 → killed by literature pre-emption

Where the idea began

R0: NGC 891 is a canonical edge-on hot halo; archive coverage appeared to leave room for a new halo-offset observation.

R2: Proposed new XMM halo mapping.

R1/R4: Examined the literature and existing observations in depth.

Decisive evidence

Hodges-Kluck, Bregman & Li 2018 already carried out the same measurement: an in-depth XMM/Chandra decomposition of the NGC 891 hot halo.

×

Not a “similar paper,” but the same measurement already published.

Project rule: a published paper that executes the same measurement → kill; it cannot be revived by changing the wording.
10 · CASE STUDY CA killed sub-claim can sometimes become part of a stronger synthesis.

I038 → I037: From single-target templates to a cross-target synthesis with a shared discriminator

I033NGC 3628
I034NGC 4945
I035NGC 4565
I036NGC 3556
I038NGC 5775standalone killed
I037 · SEED

Conditional five-target thermal-complexity comparison

Retain I038’s surviving fifth target cell, but use Chandra-defined masks, independent (not joint) fits, preregistered nested models, a model-selection threshold, and a net-count floor.

Lesson: the portfolio judge does more than “keep / delete.” It can absorb duplicate templates, retain the genuine increment, and elevate it into a comparable population question.
11 · BEST IDEAThe current leader is not finished; it is the residual question that has best survived scrutiny.

I002: NGC 253 Two-Sided Superwind · SE Halo Offset

DEVELOPING · ELO 1223.6

Question: In mirror-matched sectors, can the SE/NW thermal-pressure ratio be distinguished from 1?

Observation: A ~110 ks XMM SE offset, paired with the existing NW 0723220101 (109.5 ks).

Scientific value: Pressure P = nekT drives wind dynamics; symmetry versus asymmetry distinguishes whether disk–halo feedback is uniform or directionally injected.

min-gate
GOODdecisiveness
ADEQUATEtier
ADEQUATEcost
GOODuniqueness
Cycle 2R4: 78° projection + background → REFRAME / EMPIRICAL; retain the narrow sub-claim.
Cycle 4R4: CO(3-2) source error → FATAL to the evidence package, not the idea; Bauer 2008 pre-emption narrows the claim to matched-sector pressure.
Cycle 5Freeze k=1.96, Amin=0.20, log-ratio CI, and the aperture protocol; decisiveness adequate → good.
12 · HEAVY LANEThe real bottleneck is not another agent that can “think.”

Separate the reasoning conveyor from data reduction and simulation

Slot AFAST CONVEYOR

LIGHT / MEDIUM shifts: M0/R0/R1/R2/R4/R5/R6-lite/L

  • Only ledger writer
  • Maintains idea files, INDEX, TRUST_TABLE, and CYCLE_LOG
  • Ordered station manifest; stop on failure
result inbox
+ hash
Slot BHEAVY DRAINER

Actual archive extraction, SAS/ESAS, R6-full, simulation, and σsys

  • Pulls tasks from FRONTIER_QUEUE by blocking priority
  • Writes only to .runtime/, the Virgo workspace, and the result inbox
  • Never directly modifies the lifecycle ledger
Operational trigger: the fast lane is healthy, but 7 HEAVY items are blocked; I002’s σsys and I005’s critical gate are both in the queue. A second reasoning brain only reaches the resource wall faster.

TAKE-HOME

An agent’s value is not that it
writes faster

1

Use archive facts, literature, and gates to separate opportunities from wishes.

2

Give an independent red-team responsibility for killing bad questions, and make each kill reason a reusable constraint.

3

Make even the strongest idea retain its blocking evidence; Elo ranking does not pass a min-gate.

4

Make HEAVY execution serve only questions that can genuinely unlock a state transition.

41 live ideas5 graveyardI002 Elo 1223.60 proposal-ready — honest by design

Canonical records: pipeline/ · shared/ideas/ · TRUST_TABLE · cycles/ · cognition-os/