Merging Methodology and Infrastructure — Integrated Design & Execution Strategy
Where profiling, IR evolution, and infrastructure meet in one system — plus MVP Kill Lines
Merging Methodology and Infrastructure — Integrated Design & Execution Strategy
Hello. This is the third post in the Role-IR series.
Part 1 covered the design philosophy — treating AI agents as execution contracts rather than prompt strings. Part 2 designed the enterprise infrastructure to actually run that philosophy.
But when I placed the two documents side by side, the connections were missing.
- Where does the profiling pipeline run within the infrastructure components?
- How does the IR self-evolution loop attach to the Instruction Control Plane?
- Is Part 1's "adapter validation" the same as Part 2's Validation Layer, or something different?
- When three feedback loops run simultaneously and conflict, who takes priority?
- Starting from MVP, when does profiling enter the picture?
This post answers those questions. It brings methodology (Part 1) and infrastructure (Part 2) together within the same system components, and defines execution strategy and Kill Lines.
Integrated Component Map: Where Methodology Sits in Infrastructure
Mapping both documents onto a single view:
┌──────────────────────────────────────────────────┐
│ Human Domain │
│ role.md (purpose/intent) │
└─────────────────────┬────────────────────────────┘
▼
┌──────────────────────────────────────────────────┐
│ Instruction Control Plane (infrastructure) │
│ + Role Compiler (methodology) │
│ + Profiling Pipeline (methodology) │
│ │
│ Instruction Role IR Model Profile │
│ Registry Store Store │
│ │
│ Capability Matching │
│ Matrix Engine │
└───��─────────────────┬────────────────────────────┘
▼
┌──────────────────────────────────────────────────┐
│ Execution Layer (infrastructure) │
│ │
│ Worker internal flow: │
│ Lease → Load input → Load IR + profile │
│ → [LLM Lowering] → [Adapter: pre-exec verify] │
│ → Tool/RAG → Model Serving call │
│ → [Assurance: result verify] → Save + feedback │
└─────────────────────┬────────────────────────────┘
▼
┌──────────────────────────────────────────────────┐
│ Feedback Layer (methodology + infra) │
│ │
│ Feedback Store ← adapter + assurance feedback │
│ ├→ IR Evolution Engine │
│ ├→ Profile correction │
│ └→ Capability Matrix update │
│ │
│ Observability Plane tracks everything │
└──────────────────────────────────────────────────┘
One core principle:
The Instruction Control Plane must hold all information needed to answer "how should this role be executed on this backend right now?" Workers execute using only what this Plane provides.
Instruction Control Plane: Extended
Part 2's Instruction Control Plane originally managed role.md, constraints.yaml, output_schema.json, and adapter_manifest.yaml.
Integrating Part 1's methodology adds:
| Addition | Purpose | Storage |
| Role IR Store | Evolving IR version management | DB (JSONB) or file store |
| Model Profile Store | Per-model behavioral profiles | DB |
| Capability Matrix | Per-backend feature table | DB or config file |
| Matching Engine | Profile × role requirements → fitness | Service module |
| Profiling Pipeline | Model profiling executor | Separate batch Job |
| IR Evolution Engine | Feedback-driven IR updater | Separate batch Job |
Where Methodology Enters the Worker
Part 2's worker execution flow:
Lease → Load input → Load instruction → Tool/RAG → LLM call → Validation → Save
After methodology integration:
Lease
→ Load input
→ Load instruction + Role IR + model profile ← extended
→ LLM Lowering (IR → backend call artifact) ← new
→ Adapter: pre-execution verification ← new
→ Tool/RAG execution
→ Model Serving Plane call
→ Assurance: result verification ← strengthened
→ Save results + update state
→ Record feedback ← new
Three changes. First, instruction loading expands to include Role IR and the model's profile — the worker needs to know the model's weaknesses to reflect them in lowering. Second, LLM Lowering and adapter verification are added between instruction loading and the actual call. Third, the existing Validation Layer is absorbed into a strengthened Assurance Layer that adds evidence verification, tool policy post-checks, and role-specific validation.
Validation vs. Assurance: Clarified
| Part 2's Validation/Guardrail | Part 1's Assurance Layer | |
| Scope | JSON schema, PII redaction, policy check, repair retry | Structural + evidence + tool policy + role-specific verification + severity handling + feedback collection |
Integration conclusion: Assurance Layer subsumes Validation Layer. Existing Validation becomes Assurance Stage 1 (structural verification). Additional stages stack on top. Rename "Validation/Guardrail Plane" to "Assurance Plane" and expand scope.
Feedback Store: Where It Lives
Two options for placing the methodology's feedback store in infrastructure:
A) Add tables to Task DB — assurance_records and lowering_records tables in the existing DB. Simple, MVP-appropriate.
B) Separate Analytics DB — Split operational and analytical DBs. Execution path uses only the operational DB; feedback flows asynchronously to the analytical DB.
MVP recommendation: A. Start with Task DB tables, split to B when scale demands it.
Observability and Feedback consume the same events from different angles — one for real-time monitoring, the other for long-term IR evolution and profile correction. Implementation-wise, a single event source that branches into two pipelines is the natural shape.
Three Feedback Loops in Infrastructure
Loop 1: Execution Loop (per request)
API Gateway ��� Orchestrator → Task DB → Queue → Worker → Assurance → Response
This is Part 2's existing flow unchanged. Lowering and adapter verification are added inside the Worker, but the infrastructure structure itself doesn't change.
Executor: Worker (Celery task). Trigger: every request. Latency: real-time.
Loop 2: IR Evolution Loop
Feedback Store → IR Evolution Engine → eval verification → IR Store update
Executor: separate batch Job (Celery beat / cron). Trigger: statistical (N failures accumulated), regression, periodic. Latency: hours to days.
Key: even when a new IR version is produced, already-running workflows are unaffected. Part 2's "version pinned at workflow start" principle applies to IR as well.
Loop 3: Profile Correction Loop
IR Evolution discovers model trait → Profile Store update → Matching Engine re-evaluation
Executor: byproduct of IR Evolution + separate profile update Job. Trigger: when IR Evolution determines "this isn't an IR problem, it's a model trait." Latency: days to weeks.
This is the slowest loop with the widest blast radius. A profile change can affect matching results for other roles, so profile corrections require human confirmation before application.
Loop Conflict Resolution
Three conflict scenarios and their resolution policies:
Conflict 1: IR Evolution vs. Profile Correction
Scenario: IR Evolution decides "raise tool_call_budget to 5" while profile correction reports "this model's tool calling reliability has dropped."
Resolution: Profile correction takes priority. IR evolution makes proposals based on the current profile — if the profile changes, those premises are invalidated. When profile correction is detected, the related IR evolution queue is paused until the corrected profile is finalized.
Conflict 2: IR Evolution vs. Execution Loop
Scenario: A new IR version is being generated while the execution loop repeatedly fails with the current IR.
Resolution: The execution loop follows the fallback_chain on the current IR and doesn't wait for IR evolution. Execution is real-time; IR evolution is async batch. Failures accumulate in the Feedback Store and feed the next IR Evolution cycle.
Conflict 3: Cross-Role Profile Correction Propagation
Scenario: Model A receives a domain adjustment for Role X, which affects its matching score for Role Y.
Resolution: Profile corrections are scoped to the specific domain. The base profile isn't changed. However, if the adjustment exceeds a threshold (e.g., 30%p deviation from base), it triggers a full base profile re-measurement.
Priority Order
Execution loop (user response) > Profile correction (foundation data) > IR evolution (optimization)
Execution never stops. Profile takes precedence over IR. IR evolution is only meaningful on top of the latest profile.
Profiling Pipeline in Infrastructure
Profiling runs as a separate pipeline isolated from production traffic.
Inputs: target model ID, measurement axes, question sets (or auto-generation config), training method information (if available).
Execution: send probe requests to Model Serving Plane → repeat same questions N times → collect and classify responses → compute statistical distributions per axis → discrimination analysis → improve question sets.
Outputs: Model Profile (stored in Profile Store) + profiling report (for human review).
Infrastructure requirements: dedicated profiling instance or off-peak execution (must not affect production traffic), cost tracking in Observability Plane, versioned storage in Profile Store for comparison with previous profiles.
DB Schema Extensions
New tables added to Part 2's existing schema:
model_profiles(profile_id PK, model_id, profiling_version,
total_probes, dimensions JSONB, domain_adjustments JSONB,
status, created_at)
role_ir_versions(ir_id PK, role_id, ir_version, generation,
ir_content JSONB, evolution_trigger, change_description,
eval_regression_status, status, created_at)
role_matchings(matching_id PK, role_id, model_id, profile_id FK,
fitness_score, strengths JSONB, weaknesses JSONB,
disqualified BOOLEAN, created_at)
assurance_records(record_id PK, task_id FK, role_id, backend,
model_id, ir_version, structural_result, evidence_result,
tool_policy_result, role_specific_result JSONB,
overall_result, severity, action_taken, created_at)
backend_capabilities(backend_id PK, backend_type,
structured_outputs BOOLEAN, tool_calling BOOLEAN,
adapter_attach BOOLEAN, guided_decoding BOOLEAN,
capability_version, last_verified_at, status)
Existing tasks table gets ir_version, profile_id, and lowering_cache_hit columns for full traceability.
Version Pinning Extended
Part 2's principle: "pin instruction_version at workflow start."
Extended: pin instruction_version + ir_version + profile_id at workflow start. This ensures all tasks within a workflow execute under consistent criteria.
Phased Rollout (Methodology Focus)
| Phase | Duration | Methodology Elements | Infrastructure Elements |
| A | 4 weeks | None | DB + Queue + Worker + Validation |
| B | 6-10 weeks | Assurance, Feedback collection, manual IR | + adapter, RAG |
| C | 3-6 months | Profiling, Matching, Lowering, IR evolution | + multi-model, autoscaling |
| D | Beyond | Automation expansion, loop coordination | + cross-backend, compliance |
Phase A: no methodology elements. But add ir_version and profile_id columns as nullable to avoid future migrations.
Phase B: Validation becomes Assurance, feedback recording begins, 1-2 roles get manually written IR. No profiling yet — but feedback starts accumulating. This data fuels Phase C.
Phase C: profiling pipeline, matching engine, LLM lowering, and IR evolution all enter. This is where methodology and infrastructure fully merge.
MVP Execution Strategy: Kill Lines Make It an Experiment, Not Just a Design
This might be the most important section.
Two Failure Patterns to Avoid
- Scope grows until implementation becomes impossible
- Conversely, killing the project without validating the core hypothesis
To avoid both: don't build the entire system at once. Compose a single verifiable flow first.
Core Validation Items
This architecture needs to prove one of three things:
- The role.md → IR conversion approach is valid
- Model profile-based lowering produces real quality differences
- Separating assurance provides operational benefit over added complexity
Proving just one gives this project independent value.
Kill Lines
Pre-defined conditions for when to stop:
- Kill Line A: Role IR doesn't produce stability beyond a simple prompt template
- Kill Line B: Profile-based lowering doesn't meaningfully improve over baseline
- Kill Line C: Assurance separation adds complexity without operational benefit
- Pivot Line: Drop full orchestration, keep only "profile-guided structured extraction layer"
The goal isn't "can we build this?" but "what, if unproven, means we should discard this idea?"
Public/Private Strategy
Not "publish everything = honesty" but "well-edited publication = strategy."
Publish: design philosophy, architecture overview, MVP code, actual results. These demonstrate technical capability.
Keep private: full IR schema internals, matching algorithm details, profiling methodology details, feedback aggregation logic. These are core IP.
Integrated Architecture Diagram
Closing
What this document connected:
- Profiling pipeline sits inside the Instruction Control Plane as Profile Store, executed as a separate batch Job
- IR self-evolution runs as a separate batch Job, following the "version pinned at workflow start" principle
- Adapter (pre-execution verification) slots between lowering and the actual call inside the Worker
- Assurance subsumes and extends the existing Validation Layer
- Loop conflicts resolve with "execution > profile > IR evolution" priority
- Adoption is gradual — no methodology in MVP, feedback collection starts in Phase B, profiling enters in Phase C
In one sentence:
Methodology defines "what to measure and how to evolve." Infrastructure defines "where and how to run it." This document defines where the two meet within the same components of the same system.
And all of it has Kill Lines attached. If it can't be proven, it gets discarded. Writing this made me realize more and more that this is closer to an experiment plan than a design document.
Thanks for reading.

