Skip to main content

Command Palette

Search for a command to run...

Merging Methodology and Infrastructure — Integrated Design & Execution Strategy

Where profiling, IR evolution, and infrastructure meet in one system — plus MVP Kill Lines

Published
•10 min read•View as Markdown
E
I build data and AI systems that have to survive real constraints: time, cost, memory, and messy integration boundaries.

Merging Methodology and Infrastructure — Integrated Design & Execution Strategy

Hello. This is the third post in the Role-IR series.

Part 1 covered the design philosophy — treating AI agents as execution contracts rather than prompt strings. Part 2 designed the enterprise infrastructure to actually run that philosophy.

But when I placed the two documents side by side, the connections were missing.

  • Where does the profiling pipeline run within the infrastructure components?
  • How does the IR self-evolution loop attach to the Instruction Control Plane?
  • Is Part 1's "adapter validation" the same as Part 2's Validation Layer, or something different?
  • When three feedback loops run simultaneously and conflict, who takes priority?
  • Starting from MVP, when does profiling enter the picture?

This post answers those questions. It brings methodology (Part 1) and infrastructure (Part 2) together within the same system components, and defines execution strategy and Kill Lines.


Integrated Component Map: Where Methodology Sits in Infrastructure

Mapping both documents onto a single view:

┌──────────────────────────────────────────────────┐
│                 Human Domain                      │
│   role.md (purpose/intent)                        │
└─────────────────────┬────────────────────────────┘
                      ▼
┌──────────────────────────────────────────────────┐
│     Instruction Control Plane (infrastructure)    │
│     + Role Compiler (methodology)                 │
│     + Profiling Pipeline (methodology)            │
│                                                  │
│   Instruction    Role IR     Model Profile        │
│   Registry       Store       Store                │
│                                                  │
│   Capability     Matching                         │
│   Matrix         Engine                           │
└───��─────────────────┬────────────────────────────┘
                      ▼
┌──────────────────────────────────────────────────┐
│              Execution Layer (infrastructure)      │
│                                                  │
│  Worker internal flow:                            │
│  Lease → Load input → Load IR + profile           │
│  → [LLM Lowering] → [Adapter: pre-exec verify]   │
│  → Tool/RAG → Model Serving call                  │
│  → [Assurance: result verify] → Save + feedback   │
└─────────────────────┬────────────────────────────┘
                      ▼
┌──────────────────────────────────────────────────┐
│         Feedback Layer (methodology + infra)       │
│                                                  │
│  Feedback Store ← adapter + assurance feedback    │
│       ├→ IR Evolution Engine                      │
│       ├→ Profile correction                       │
│       └→ Capability Matrix update                 │
│                                                  │
│  Observability Plane tracks everything            │
└──────────────────────────────────────────────────┘

One core principle:

The Instruction Control Plane must hold all information needed to answer "how should this role be executed on this backend right now?" Workers execute using only what this Plane provides.


Instruction Control Plane: Extended

Part 2's Instruction Control Plane originally managed role.md, constraints.yaml, output_schema.json, and adapter_manifest.yaml.

Integrating Part 1's methodology adds:

AdditionPurposeStorage
Role IR StoreEvolving IR version managementDB (JSONB) or file store
Model Profile StorePer-model behavioral profilesDB
Capability MatrixPer-backend feature tableDB or config file
Matching EngineProfile × role requirements → fitnessService module
Profiling PipelineModel profiling executorSeparate batch Job
IR Evolution EngineFeedback-driven IR updaterSeparate batch Job

Where Methodology Enters the Worker

Part 2's worker execution flow:

Lease → Load input → Load instruction → Tool/RAG → LLM call → Validation → Save

After methodology integration:

Lease
  → Load input
  → Load instruction + Role IR + model profile    ← extended
  → LLM Lowering (IR → backend call artifact)      ← new
  → Adapter: pre-execution verification             ← new
  → Tool/RAG execution
  → Model Serving Plane call
  → Assurance: result verification                  ← strengthened
  → Save results + update state
  → Record feedback                                 ← new

Three changes. First, instruction loading expands to include Role IR and the model's profile — the worker needs to know the model's weaknesses to reflect them in lowering. Second, LLM Lowering and adapter verification are added between instruction loading and the actual call. Third, the existing Validation Layer is absorbed into a strengthened Assurance Layer that adds evidence verification, tool policy post-checks, and role-specific validation.


Validation vs. Assurance: Clarified

Part 2's Validation/GuardrailPart 1's Assurance Layer
ScopeJSON schema, PII redaction, policy check, repair retryStructural + evidence + tool policy + role-specific verification + severity handling + feedback collection

Integration conclusion: Assurance Layer subsumes Validation Layer. Existing Validation becomes Assurance Stage 1 (structural verification). Additional stages stack on top. Rename "Validation/Guardrail Plane" to "Assurance Plane" and expand scope.


Feedback Store: Where It Lives

Two options for placing the methodology's feedback store in infrastructure:

A) Add tables to Task DB — assurance_records and lowering_records tables in the existing DB. Simple, MVP-appropriate.

B) Separate Analytics DB — Split operational and analytical DBs. Execution path uses only the operational DB; feedback flows asynchronously to the analytical DB.

MVP recommendation: A. Start with Task DB tables, split to B when scale demands it.

Observability and Feedback consume the same events from different angles — one for real-time monitoring, the other for long-term IR evolution and profile correction. Implementation-wise, a single event source that branches into two pipelines is the natural shape.


Three Feedback Loops in Infrastructure

Loop 1: Execution Loop (per request)

API Gateway ��� Orchestrator → Task DB → Queue → Worker → Assurance → Response

This is Part 2's existing flow unchanged. Lowering and adapter verification are added inside the Worker, but the infrastructure structure itself doesn't change.

Executor: Worker (Celery task). Trigger: every request. Latency: real-time.

Loop 2: IR Evolution Loop

Feedback Store → IR Evolution Engine → eval verification → IR Store update

Executor: separate batch Job (Celery beat / cron). Trigger: statistical (N failures accumulated), regression, periodic. Latency: hours to days.

Key: even when a new IR version is produced, already-running workflows are unaffected. Part 2's "version pinned at workflow start" principle applies to IR as well.

Loop 3: Profile Correction Loop

IR Evolution discovers model trait → Profile Store update → Matching Engine re-evaluation

Executor: byproduct of IR Evolution + separate profile update Job. Trigger: when IR Evolution determines "this isn't an IR problem, it's a model trait." Latency: days to weeks.

This is the slowest loop with the widest blast radius. A profile change can affect matching results for other roles, so profile corrections require human confirmation before application.


Loop Conflict Resolution

Three conflict scenarios and their resolution policies:

Conflict 1: IR Evolution vs. Profile Correction

Scenario: IR Evolution decides "raise tool_call_budget to 5" while profile correction reports "this model's tool calling reliability has dropped."

Resolution: Profile correction takes priority. IR evolution makes proposals based on the current profile — if the profile changes, those premises are invalidated. When profile correction is detected, the related IR evolution queue is paused until the corrected profile is finalized.

Conflict 2: IR Evolution vs. Execution Loop

Scenario: A new IR version is being generated while the execution loop repeatedly fails with the current IR.

Resolution: The execution loop follows the fallback_chain on the current IR and doesn't wait for IR evolution. Execution is real-time; IR evolution is async batch. Failures accumulate in the Feedback Store and feed the next IR Evolution cycle.

Conflict 3: Cross-Role Profile Correction Propagation

Scenario: Model A receives a domain adjustment for Role X, which affects its matching score for Role Y.

Resolution: Profile corrections are scoped to the specific domain. The base profile isn't changed. However, if the adjustment exceeds a threshold (e.g., 30%p deviation from base), it triggers a full base profile re-measurement.

Priority Order

Execution loop (user response) > Profile correction (foundation data) > IR evolution (optimization)

Execution never stops. Profile takes precedence over IR. IR evolution is only meaningful on top of the latest profile.


Profiling Pipeline in Infrastructure

Profiling runs as a separate pipeline isolated from production traffic.

Inputs: target model ID, measurement axes, question sets (or auto-generation config), training method information (if available).

Execution: send probe requests to Model Serving Plane → repeat same questions N times → collect and classify responses → compute statistical distributions per axis → discrimination analysis → improve question sets.

Outputs: Model Profile (stored in Profile Store) + profiling report (for human review).

Infrastructure requirements: dedicated profiling instance or off-peak execution (must not affect production traffic), cost tracking in Observability Plane, versioned storage in Profile Store for comparison with previous profiles.


DB Schema Extensions

New tables added to Part 2's existing schema:

model_profiles(profile_id PK, model_id, profiling_version,
  total_probes, dimensions JSONB, domain_adjustments JSONB,
  status, created_at)

role_ir_versions(ir_id PK, role_id, ir_version, generation,
  ir_content JSONB, evolution_trigger, change_description,
  eval_regression_status, status, created_at)

role_matchings(matching_id PK, role_id, model_id, profile_id FK,
  fitness_score, strengths JSONB, weaknesses JSONB,
  disqualified BOOLEAN, created_at)

assurance_records(record_id PK, task_id FK, role_id, backend,
  model_id, ir_version, structural_result, evidence_result,
  tool_policy_result, role_specific_result JSONB,
  overall_result, severity, action_taken, created_at)

backend_capabilities(backend_id PK, backend_type,
  structured_outputs BOOLEAN, tool_calling BOOLEAN,
  adapter_attach BOOLEAN, guided_decoding BOOLEAN,
  capability_version, last_verified_at, status)

Existing tasks table gets ir_version, profile_id, and lowering_cache_hit columns for full traceability.

Version Pinning Extended

Part 2's principle: "pin instruction_version at workflow start."

Extended: pin instruction_version + ir_version + profile_id at workflow start. This ensures all tasks within a workflow execute under consistent criteria.


Phased Rollout (Methodology Focus)

PhaseDurationMethodology ElementsInfrastructure Elements
A4 weeksNoneDB + Queue + Worker + Validation
B6-10 weeksAssurance, Feedback collection, manual IR+ adapter, RAG
C3-6 monthsProfiling, Matching, Lowering, IR evolution+ multi-model, autoscaling
DBeyondAutomation expansion, loop coordination+ cross-backend, compliance

Phase A: no methodology elements. But add ir_version and profile_id columns as nullable to avoid future migrations.

Phase B: Validation becomes Assurance, feedback recording begins, 1-2 roles get manually written IR. No profiling yet — but feedback starts accumulating. This data fuels Phase C.

Phase C: profiling pipeline, matching engine, LLM lowering, and IR evolution all enter. This is where methodology and infrastructure fully merge.


MVP Execution Strategy: Kill Lines Make It an Experiment, Not Just a Design

This might be the most important section.

Two Failure Patterns to Avoid

  • Scope grows until implementation becomes impossible
  • Conversely, killing the project without validating the core hypothesis

To avoid both: don't build the entire system at once. Compose a single verifiable flow first.

Core Validation Items

This architecture needs to prove one of three things:

  1. The role.md → IR conversion approach is valid
  2. Model profile-based lowering produces real quality differences
  3. Separating assurance provides operational benefit over added complexity

Proving just one gives this project independent value.

Kill Lines

Pre-defined conditions for when to stop:

  • Kill Line A: Role IR doesn't produce stability beyond a simple prompt template
  • Kill Line B: Profile-based lowering doesn't meaningfully improve over baseline
  • Kill Line C: Assurance separation adds complexity without operational benefit
  • Pivot Line: Drop full orchestration, keep only "profile-guided structured extraction layer"

The goal isn't "can we build this?" but "what, if unproven, means we should discard this idea?"

Public/Private Strategy

Not "publish everything = honesty" but "well-edited publication = strategy."

Publish: design philosophy, architecture overview, MVP code, actual results. These demonstrate technical capability.

Keep private: full IR schema internals, matching algorithm details, profiling methodology details, feedback aggregation logic. These are core IP.


Integrated Architecture Diagram


Closing

What this document connected:

  1. Profiling pipeline sits inside the Instruction Control Plane as Profile Store, executed as a separate batch Job
  2. IR self-evolution runs as a separate batch Job, following the "version pinned at workflow start" principle
  3. Adapter (pre-execution verification) slots between lowering and the actual call inside the Worker
  4. Assurance subsumes and extends the existing Validation Layer
  5. Loop conflicts resolve with "execution > profile > IR evolution" priority
  6. Adoption is gradual — no methodology in MVP, feedback collection starts in Phase B, profiling enters in Phase C

In one sentence:

Methodology defines "what to measure and how to evolve." Infrastructure defines "where and how to run it." This document defines where the two meet within the same components of the same system.

And all of it has Kill Lines attached. If it can't be proven, it gets discarded. Writing this made me realize more and more that this is closer to an experiment plan than a design document.

Thanks for reading.