live
MODEL IDanthropic-claude-opus-4-7

Claude Opus 4.7

Anthropic
by 1M context

Claude Opus 4.7 is Anthropic's highest-capability generally available Claude model. It is designed for demanding coding, agent orchestration, multimodal reasoning, and high-stakes professional workflows, with stronger instruction following, better high-resolution vision, adaptive thinking, and a 1M-token context window.

Claude Opus 4.7

reference-to-text

Beverage Shelf Compliance Audit$0.0609~36s
# Refrigerated Beverage Bay – Merchandising Compliance Report

## 1. Executive Summary
The cooler bay is broadly well-stocked with strong vertical variety, transitioning from energy drinks and canned sodas at the top to sports/tea bottles and waters at the bottom. However, the middle-right section (Shelf 4 center-right and Shelf 5 center) shows the most commercially significant issue: a visible empty wire zone and only two isolated bottles standing alone, creating a noticeable "hole" in the customer sightline. Several shelf-edge labels appear present but are illegible in the image, and at least two bright yellow promotional tags on Shelf 4 suggest active price callouts that cannot be verified. Overall facing discipline is acceptable, but front-facing alignment and gap-filling on Shelves 4–5 should be prioritized.

## 2. Shelf Inventory Table
Numbered top (1) to bottom (6). Facing counts are visual estimates; ranges reflect blur and overlap.

| # | Visible Category / Family | Occupied Facings (est.) | Empty Facings (est.) | Notable Placement Issues | Confidence |
|---|---|---|---|---|---|
| 1 | Energy drink cans (dark/green left) + assorted slim cans (center/right), silver/red cans (far right) | ~18–22 | 0 visible | Full, tidy front row | Medium |
| 2 | Energy/soda cans, pink, blue, silver, yellow-green cans | ~20–24 | 0–1 | Dense; some cans behind front row appear staggered | Medium |
| 3 | Green/purple energy cans (left), small yellow cans (center), amber PET bottles (right) | ~22–26 | 0–2 | Category break mid-shelf appears intentional | Medium |
| 4 | Black & blue PET bottles (left), red/orange cola-style PET, yellow bottle, clear water bottles (right) | ~16–20 | 2–4 | Single yellow bottle isolated center; possible gap right of center | Low–Medium |
| 5 | Orange/red sports or tea bottles (left), **largely empty center**, water bottles (right) | ~10–14 | 6–10 | **Major void in center**; only 1 gold can + 1 brown bottle standing alone | High (for the gap) |
| 6 | Green-capped PET (left), amber PET (center-left), red/orange PET (center), clear waters (right) | ~22–28 | 0–2 | Full front row; good facing | Medium |

## 3. Exceptions Log

- **Shelf 5, center:** Significant stockout / low-stock zone. Only two items visible (one gold/yellow can and one brown bottle) on an otherwise empty wire shelf. **Highest-priority gap.**
- **Shelf 4, center:** A lone yellow bottle sits between the red/orange bottle block and the water block, suggesting either a misplaced SKU or a nearly depleted facing.
- **Shelf 4, left of center:** Two yellow promotional tags mounted on the shelf edge — legibility poor; cannot confirm they correspond to the products directly above.
- **Shelf 6, far left:** One green-capped bottle appears shorter/differently shaped than its neighbors — possible misplaced SKU or mixed pack; low confidence.
- **Shelves 1–3:** Shelf-edge label strips are visible but text is not legible in the image; cannot confirm label-to-product alignment.
- **Shelf 5, right block (waters):** Front row appears complete, but back rows look sparse behind the front facings — potential replenishment need, low confidence.
- No fallen products, no visibly blocked primary labels beyond general image blur.

## 4. Pricing and Promotion Review
Legible text transcription is extremely limited due to image resolution.

- Shelf-edge strips on all six shelves: **not legible**.
- Shelf 4 promotional tags (yellow, two adjacent): **not legible** — color and placement suggest price/promo callouts, but no digits or words can be reliably read.
- Shelf 6 far-left small yellow tag: **not legible**.
- No multi-buy, percentage-off, or currency figures can be confirmed.

## 5. Priority Actions

**Immediate fixes (during current shift):**
1. Fill the center void on **Shelf 5** — either replenish the intended SKU or pull-forward/consolidate adjacent product to eliminate the visible hole.
2. Resolve the isolated yellow bottle on **Shelf 4, center** — return to correct home if misplaced, or front-face the remaining stock.
3. Verify the two yellow promo tags on **Shelf 4** correspond to the products directly above and that pricing is current.

**Replenishment follow-up (next stock cycle):**
4. Audit back-row depth on **Shelf 5 waters (right)** and **Shelf 4 waters (right)** to confirm sufficient reserve behind the front row.
5. Confirm/reprint shelf-edge labels on **Shelves 1–3** so category and price are clearly readable to shoppers.
6. Spot-check the anomalous green-capped bottle on **Shelf 6, far left** for SKU accuracy.

## 6. Evidence Limitations
- **Resolution & focus:** Label text on cans, bottles, and shelf strips is uniformly soft; brand names are not reliably readable, so no brand identification is offered.
- **Glare & reflection:** Overhead LED strips and the glass door create highlights across the top of each shelf, which may mask small labels or additional promo tags.
- **Perspective:** The image is close to straight-on but slightly elevated, so back-row depth on lower shelves is partially hidden by front facings — stockouts behind the first row cannot be confirmed.
- **Occlusion:** Front-row bottles conceal secondary rows, particularly on Shelves 4–6.
- **No side context:** Adjacent bays, floor stock, or planogram documentation are not visible; compliance to any specific planogram cannot be assessed.
Finish reason: stop
Try in Playground
import { createClient } from '@runware/sdk'

const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()

const [result] = await client.run({
  model: 'anthropic:claude@opus-4.7',
  settings: {
    thinkingLevel: 'adaptive'
  },
  inputs: {
    images: [
      'https://assets.runware.ai/assets/inputs/41ceb2bf-607f-464a-b09e-34b00d03b09a.jpg'
    ]
  },
  messages: [
    {
      role: 'user',
      content: 'Act as a retail merchandising auditor. Inspect the attached photograph of the refrigerated beverage bay and produce a concise, field-ready compliance report grounded only in details that are visibly supported by the image.\n\nStructure the report as follows:\n\n1. Executive summary: 3–5 sentences describing the display\'s overall condition and the most commercially significant issues.\n2. Shelf inventory table: one row per shelf, numbered top to bottom. Include visible product category or recognizable product family, estimated occupied facings, estimated empty facings, notable placement issues, and confidence level.\n3. Exceptions log: list each visible stockout, low-stock row, misplaced item, fallen or poorly faced product, blocked label, missing or mismatched shelf label, and promotional inconsistency. For each exception, identify its location precisely using shelf number and left/center/right position.\n4. Pricing and promotion review: transcribe only shelf-label or promotional text that is genuinely legible. Place uncertain characters in square brackets and write “not legible” rather than guessing.\n5. Priority actions: provide a ranked list of no more than six corrective actions for the store team, separating immediate fixes from replenishment follow-up.\n6. Evidence limitations: note glare, occlusion, perspective, resolution, or other factors that prevent reliable conclusions.\n\nDo not invent brand names, prices, planogram requirements, SKU identities, or items outside the frame. Distinguish confirmed observations from estimates. Where exact facing counts are ambiguous, give a range and explain why.'
    }
  ]
})
import asyncio
import os

from runware import Runware


async def main():
    async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
        results = await client.run({
            "model": "anthropic:claude@opus-4.7",
            "settings": {
                "thinkingLevel": "adaptive"
            },
            "inputs": {
                "images": [
                    "https://assets.runware.ai/assets/inputs/41ceb2bf-607f-464a-b09e-34b00d03b09a.jpg"
                ]
            },
            "messages": [
                {
                    "role": "user",
                    "content": "Act as a retail merchandising auditor. Inspect the attached photograph of the refrigerated beverage bay and produce a concise, field-ready compliance report grounded only in details that are visibly supported by the image.\n\nStructure the report as follows:\n\n1. Executive summary: 3–5 sentences describing the display's overall condition and the most commercially significant issues.\n2. Shelf inventory table: one row per shelf, numbered top to bottom. Include visible product category or recognizable product family, estimated occupied facings, estimated empty facings, notable placement issues, and confidence level.\n3. Exceptions log: list each visible stockout, low-stock row, misplaced item, fallen or poorly faced product, blocked label, missing or mismatched shelf label, and promotional inconsistency. For each exception, identify its location precisely using shelf number and left/center/right position.\n4. Pricing and promotion review: transcribe only shelf-label or promotional text that is genuinely legible. Place uncertain characters in square brackets and write “not legible” rather than guessing.\n5. Priority actions: provide a ranked list of no more than six corrective actions for the store team, separating immediate fixes from replenishment follow-up.\n6. Evidence limitations: note glare, occlusion, perspective, resolution, or other factors that prevent reliable conclusions.\n\nDo not invent brand names, prices, planogram requirements, SKU identities, or items outside the frame. Distinguish confirmed observations from estimates. Where exact facing counts are ambiguous, give a range and explain why."
                }
            ]
        })


asyncio.run(main())
curl https://api.runware.ai/v1 \
  -H "Authorization: Bearer $RUNWARE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '[
    {
      "taskType": "textInference",
      "taskUUID": "fa0a864a-0e73-418a-a7eb-e9a4cb8fff7c",
      "model": "anthropic:claude@opus-4.7",
      "settings": {
        "thinkingLevel": "adaptive"
      },
      "inputs": {
        "images": [
          "https://assets.runware.ai/assets/inputs/41ceb2bf-607f-464a-b09e-34b00d03b09a.jpg"
        ]
      },
      "messages": [
        {
          "role": "user",
          "content": "Act as a retail merchandising auditor. Inspect the attached photograph of the refrigerated beverage bay and produce a concise, field-ready compliance report grounded only in details that are visibly supported by the image.\n\nStructure the report as follows:\n\n1. Executive summary: 3–5 sentences describing the display's overall condition and the most commercially significant issues.\n2. Shelf inventory table: one row per shelf, numbered top to bottom. Include visible product category or recognizable product family, estimated occupied facings, estimated empty facings, notable placement issues, and confidence level.\n3. Exceptions log: list each visible stockout, low-stock row, misplaced item, fallen or poorly faced product, blocked label, missing or mismatched shelf label, and promotional inconsistency. For each exception, identify its location precisely using shelf number and left/center/right position.\n4. Pricing and promotion review: transcribe only shelf-label or promotional text that is genuinely legible. Place uncertain characters in square brackets and write “not legible” rather than guessing.\n5. Priority actions: provide a ranked list of no more than six corrective actions for the store team, separating immediate fixes from replenishment follow-up.\n6. Evidence limitations: note glare, occlusion, perspective, resolution, or other factors that prevent reliable conclusions.\n\nDo not invent brand names, prices, planogram requirements, SKU identities, or items outside the frame. Distinguish confirmed observations from estimates. Where exact facing counts are ambiguous, give a range and explain why."
        }
      ]
    }
  ]'
runware run anthropic:claude@opus-4.7 \
  settings.thinkingLevel=adaptive \
  inputs.images.0=https://assets.runware.ai/assets/inputs/41ceb2bf-607f-464a-b09e-34b00d03b09a.jpg \
  messages.0.role=user \
  messages.0.content="Act as a retail merchandising auditor. Inspect the attached photograph of the refrigerated beverage bay and produce a concise, field-ready compliance report grounded only in details that are visibly supported by the image.

Structure the report as follows:

1. Executive summary: 3–5 sentences describing the display's overall condition and the most commercially significant issues.
2. Shelf inventory table: one row per shelf, numbered top to bottom. Include visible product category or recognizable product family, estimated occupied facings, estimated empty facings, notable placement issues, and confidence level.
3. Exceptions log: list each visible stockout, low-stock row, misplaced item, fallen or poorly faced product, blocked label, missing or mismatched shelf label, and promotional inconsistency. For each exception, identify its location precisely using shelf number and left/center/right position.
4. Pricing and promotion review: transcribe only shelf-label or promotional text that is genuinely legible. Place uncertain characters in square brackets and write “not legible” rather than guessing.
5. Priority actions: provide a ranked list of no more than six corrective actions for the store team, separating immediate fixes from replenishment follow-up.
6. Evidence limitations: note glare, occlusion, perspective, resolution, or other factors that prevent reliable conclusions.

Do not invent brand names, prices, planogram requirements, SKU identities, or items outside the frame. Distinguish confirmed observations from estimates. Where exact facing counts are ambiguous, give a range and explain why."
{
  "taskType": "textInference",
  "taskUUID": "fa0a864a-0e73-418a-a7eb-e9a4cb8fff7c",
  "model": "anthropic:claude@opus-4.7",
  "settings": {
    "thinkingLevel": "adaptive"
  },
  "inputs": {
    "images": [
      "https://assets.runware.ai/assets/inputs/41ceb2bf-607f-464a-b09e-34b00d03b09a.jpg"
    ]
  },
  "messages": [
    {
      "role": "user",
      "content": "Act as a retail merchandising auditor. Inspect the attached photograph of the refrigerated beverage bay and produce a concise, field-ready compliance report grounded only in details that are visibly supported by the image.\n\nStructure the report as follows:\n\n1. Executive summary: 3–5 sentences describing the display's overall condition and the most commercially significant issues.\n2. Shelf inventory table: one row per shelf, numbered top to bottom. Include visible product category or recognizable product family, estimated occupied facings, estimated empty facings, notable placement issues, and confidence level.\n3. Exceptions log: list each visible stockout, low-stock row, misplaced item, fallen or poorly faced product, blocked label, missing or mismatched shelf label, and promotional inconsistency. For each exception, identify its location precisely using shelf number and left/center/right position.\n4. Pricing and promotion review: transcribe only shelf-label or promotional text that is genuinely legible. Place uncertain characters in square brackets and write “not legible” rather than guessing.\n5. Priority actions: provide a ranked list of no more than six corrective actions for the store team, separating immediate fixes from replenishment follow-up.\n6. Evidence limitations: note glare, occlusion, perspective, resolution, or other factors that prevent reliable conclusions.\n\nDo not invent brand names, prices, planogram requirements, SKU identities, or items outside the frame. Distinguish confirmed observations from estimates. Where exact facing counts are ambiguous, give a range and explain why."
    }
  ]
}
Response
{
  "taskType": "textInference",
  "taskUUID": "fa0a864a-0e73-418a-a7eb-e9a4cb8fff7c",
  "text": "# Refrigerated Beverage Bay – Merchandising Compliance Report\n\n## 1. Executive Summary\nThe cooler bay is broadly well-stocked with strong vertical variety, transitioning from energy drinks and canned sodas at the top to sports/tea bottles and waters at the bottom. However, the middle-right section (Shelf 4 center-right and Shelf 5 center) shows the most commercially significant issue: a visible empty wire zone and only two isolated bottles standing alone, creating a noticeable \"hole\" in the customer sightline. Several shelf-edge labels appear present but are illegible in the image, and at least two bright yellow promotional tags on Shelf 4 suggest active price callouts that cannot be verified. Overall facing discipline is acceptable, but front-facing alignment and gap-filling on Shelves 4–5 should be prioritized.\n\n## 2. Shelf Inventory Table\nNumbered top (1) to bottom (6). Facing counts are visual estimates; ranges reflect blur and overlap.\n\n| # | Visible Category / Family | Occupied Facings (est.) | Empty Facings (est.) | Notable Placement Issues | Confidence |\n|---|---|---|---|---|---|\n| 1 | Energy drink cans (dark/green left) + assorted slim cans (center/right), silver/red cans (far right) | ~18–22 | 0 visible | Full, tidy front row | Medium |\n| 2 | Energy/soda cans, pink, blue, silver, yellow-green cans | ~20–24 | 0–1 | Dense; some cans behind front row appear staggered | Medium |\n| 3 | Green/purple energy cans (left), small yellow cans (center), amber PET bottles (right) | ~22–26 | 0–2 | Category break mid-shelf appears intentional | Medium |\n| 4 | Black & blue PET bottles (left), red/orange cola-style PET, yellow bottle, clear water bottles (right) | ~16–20 | 2–4 | Single yellow bottle isolated center; possible gap right of center | Low–Medium |\n| 5 | Orange/red sports or tea bottles (left), **largely empty center**, water bottles (right) | ~10–14 | 6–10 | **Major void in center**; only 1 gold can + 1 brown bottle standing alone | High (for the gap) |\n| 6 | Green-capped PET (left), amber PET (center-left), red/orange PET (center), clear waters (right) | ~22–28 | 0–2 | Full front row; good facing | Medium |\n\n## 3. Exceptions Log\n\n- **Shelf 5, center:** Significant stockout / low-stock zone. Only two items visible (one gold/yellow can and one brown bottle) on an otherwise empty wire shelf. **Highest-priority gap.**\n- **Shelf 4, center:** A lone yellow bottle sits between the red/orange bottle block and the water block, suggesting either a misplaced SKU or a nearly depleted facing.\n- **Shelf 4, left of center:** Two yellow promotional tags mounted on the shelf edge — legibility poor; cannot confirm they correspond to the products directly above.\n- **Shelf 6, far left:** One green-capped bottle appears shorter/differently shaped than its neighbors — possible misplaced SKU or mixed pack; low confidence.\n- **Shelves 1–3:** Shelf-edge label strips are visible but text is not legible in the image; cannot confirm label-to-product alignment.\n- **Shelf 5, right block (waters):** Front row appears complete, but back rows look sparse behind the front facings — potential replenishment need, low confidence.\n- No fallen products, no visibly blocked primary labels beyond general image blur.\n\n## 4. Pricing and Promotion Review\nLegible text transcription is extremely limited due to image resolution.\n\n- Shelf-edge strips on all six shelves: **not legible**.\n- Shelf 4 promotional tags (yellow, two adjacent): **not legible** — color and placement suggest price/promo callouts, but no digits or words can be reliably read.\n- Shelf 6 far-left small yellow tag: **not legible**.\n- No multi-buy, percentage-off, or currency figures can be confirmed.\n\n## 5. Priority Actions\n\n**Immediate fixes (during current shift):**\n1. Fill the center void on **Shelf 5** — either replenish the intended SKU or pull-forward/consolidate adjacent product to eliminate the visible hole.\n2. Resolve the isolated yellow bottle on **Shelf 4, center** — return to correct home if misplaced, or front-face the remaining stock.\n3. Verify the two yellow promo tags on **Shelf 4** correspond to the products directly above and that pricing is current.\n\n**Replenishment follow-up (next stock cycle):**\n4. Audit back-row depth on **Shelf 5 waters (right)** and **Shelf 4 waters (right)** to confirm sufficient reserve behind the front row.\n5. Confirm/reprint shelf-edge labels on **Shelves 1–3** so category and price are clearly readable to shoppers.\n6. Spot-check the anomalous green-capped bottle on **Shelf 6, far left** for SKU accuracy.\n\n## 6. Evidence Limitations\n- **Resolution & focus:** Label text on cans, bottles, and shelf strips is uniformly soft; brand names are not reliably readable, so no brand identification is offered.\n- **Glare & reflection:** Overhead LED strips and the glass door create highlights across the top of each shelf, which may mask small labels or additional promo tags.\n- **Perspective:** The image is close to straight-on but slightly elevated, so back-row depth on lower shelves is partially hidden by front facings — stockouts behind the first row cannot be confirmed.\n- **Occlusion:** Front-row bottles conceal secondary rows, particularly on Shelves 4–6.\n- **No side context:** Adjacent bays, floor stock, or planogram documentation are not visible; compliance to any specific planogram cannot be assessed.",
  "cost": 0.06087,
  "finishReason": "stop"
}

text-generation

Enterprise Identity Consolidation Execution Plan$0.438~3m 50s
# Entra ID Tenant Consolidation: Northstar Mutual ← Harbor Specialty
**Program Execution Plan — Version 1.0 (Decision-Ready Draft)**
Prepared by: Principal Identity Architect
Audience: CIO, CISO, service owners, migration team, steering committee

---

## Feasibility Statement (Read First)

The 28-week deadline is **achievable only under a scoped-migration pattern with two conditions**:

1. The **34 LDAP applications** are **not** re-platformed inside the program window. They remain on a coexistence Active Directory footprint (either a retained Harbor AD forest trusted from Northstar, or a new resource forest) and are re-platformed on a follow-on 6–9 month track.
2. A **specialist migration vendor is engaged in Week 1–4** for tooling (cross-tenant sync accelerator, mailbox and OneDrive movement, correlation reporting). Without vendor tooling, in-house staffing is the binding constraint (see §13).

If executives insist on LDAP re-platforming inside 28 weeks, the defensible options are (a) extend to 40 weeks, (b) add three application engineers and one identity architect, or (c) descope 12 Azure subscription migrations to a follow-on wave. This document assumes the scoped pattern.

---

## 1. Executive Recommendation

### Proposed pattern: **Staged coexistence-first consolidation into Tenant N**, using cross-tenant synchronization (CTS), progressive mail routing, and wave-based user cutover with source-tenant retained read-only for 90 days.

| Element | Choice |
|---|---|
| Destination | Tenant N (cloud-native, PIM-mature, passwordless-forward) |
| Coexistence engine | Entra Cross-Tenant Synchronization (B2B members) + mail coexistence via smart host |
| Directory synchronization | Retain Harbor AD Connect during coexistence; cutover to Tenant N-anchored Connect for the two Harbor forests; groom out post-migration |
| Application handling | Re-federate SAML/OIDC to Tenant N; SCIM re-point; LDAP defer; object-ID apps handled per §6 |
| Data movement | Native cross-tenant mailbox migration + SPMT/Mover for SharePoint/OneDrive; Teams via cross-tenant Teams migration |
| Devices | Intune tenant-to-tenant using Windows Autopilot re-enrollment for Windows; ADE re-enrollment for macOS; MAM re-enrollment for mobile |
| Contractor devices | B2B guest with device compliance enforced via CA; no MDM enrollment |

### Why this over plausible alternatives

| Alternative | Why rejected |
|---|---|
| Big-bang cutover in one weekend | 4,600 employees + 900 contractors + 137 apps + 12 subscriptions exceeds recoverable blast radius; violates 15-minute SLA. |
| Keep both tenants indefinitely with B2B only | Fails consolidation mandate, does not resolve privileged sprawl, sustains SMS-MFA debt, does not reduce license spend. |
| Migrate to a **new** third tenant | Doubles work; abandons Tenant N's mature PIM, CA, and passwordless posture. |
| Rebuild Harbor users in Tenant N from HR feed and abandon source | Loses mailbox history, breaks legal holds and retention labels, breaks object-ID dependent apps beyond repair. |

### Five largest risks (full register in §12)
1. **Object-ID-anchored apps without vendor remap** — three apps with no bulk remap; per-user re-linking at 5,500 users is infeasible without downtime concessions or bridging identity mapping.
2. **Privileged-access exposure during coexistence** — 39 GAs, 11 standing, 6 shared emergency accounts in tenant H create a large attack surface until harmonized.
3. **Correlation accuracy below 99.95%** — 1,200 guest collisions, 87 duplicate proxies, 46 dual identities; correlation gate directly controls schedule.
4. **Mail routing outage during proxy-address hand-off** — customer-facing 15-minute SLA is unforgiving; misconfigured accepted domains or MX cutover causes hours of NDRs.
5. **Change-freeze compression** — quarter-end (W9–11) and catastrophe response (W18–21) remove ~7 productive weeks; correlation, pilots, and final waves must fit around them.

### Executive decisions required within 14 days

| # | Decision | Owner | Trigger if not made |
|---|---|---|---|
| D1 | Approve scoped pattern (LDAP deferred) or fund extension | CIO | 28-week plan slips |
| D2 | Approve vendor engagement and $180K envelope | CIO + CFO | Correlation and tooling delayed; slip ≥4 weeks |
| D3 | Approve retirement of 6 Harbor shared emergency accounts and mandate PIM+FIDO2 for all admins | CISO | Non-negotiable Control 2 unmet |
| D4 | Confirm Tenant N as authoritative and endorse `harborspecialty.com` addition as accepted domain | CIO + Legal | Namespace and email strategy blocked |
| D5 | Legal/Privacy officers confirm cross-border processing model (US/UK/DE) and legal-hold custodianship transfer approach | Compliance Lead + Counsel | Control 4 unmet; mailbox migration blocked |
| D6 | Approve contractor identity model (B2B members vs. guests) and device policy for 760 BYOD | CISO | Contractor cutover blocked |
| D7 | Approve freeze exceptions policy and go/no-go authority (single named role) | CIO | Waves stall in freeze weeks |

---

## 2. Current-State Problem Decomposition

| Domain | Tenant N | Tenant H | Gap / Risk |
|---|---|---|---|
| Identity authority | Cloud-first, HRIS-driven | Two on-prem AD forests → Entra Connect | Two forests must be reconciled; ownership of source-of-truth per attribute unclear |
| Authentication | 68% passwordless (FIDO2/WHfB); CA baseline | Password + SMS MFA dominant | SMS MFA is not phishing-resistant; must uplift before privileged cutover |
| Authorization | Group-based, PIM for privileged | Ad hoc; 39 GAs; 11 standing | Standing privilege violates least privilege |
| Lifecycle (JML) | HR-driven, automated | Manual + scripted per forest | JML gaps produce dormant accounts; contractors unmanaged at scale |
| Applications | Consolidated SSO catalog | 73 SAML/OIDC, 18 SCIM, 34 LDAP, 7 object-ID-dependent | LDAP apps block full decommission; object-ID apps threaten users |
| Messaging & collab | Single EXO org | Separate EXO; overlapping proxies (87 duplicates) | Proxy collisions block mail routing |
| Endpoints | Intune-managed, Autopilot | Intune, mixed enrollment maturity | Kiosks and BYOD contractors need bespoke approach |
| Azure resources | Established landing zones | 12 subscriptions, unknown RBAC hygiene | Subscription transfer requires billing + RBAC redesign |
| Privileged access | PIM, 2 monitored break-glass | 39 GA, 11 standing, 6 shared emergency | Immediate risk; cannot enter Tenant N as-is |
| Compliance | Retention labels, legal holds mapped | Unknown label parity, hold inventory needed | Control 4 blocker |
| Support readiness | Mature service desk | Separate ITSM, separate runbooks | Merged runbook + bilingual (EN/DE) support required |

**Unknowns to close in Weeks 1–3:** exact legal-hold custodian list; Harbor SCIM app tokens and rotation authority; Harbor conditional access baseline; Azure subscription commercial ownership (EA vs. MCA); DE works-council notification status.

---

## 3. Target-State Architecture

### Permanent design
- **Authoritative sources:** HRIS (Workday) → Tenant N Entra ID for employees. Contractor management system → Tenant N with sponsor attribute mandatory. Application ownership records → ServiceNow CMDB.
- **JML:** Workday event → Entra provisioning → downstream SCIM apps → Exchange, Teams, SharePoint, Intune. Leaver: 24-hour disable, 30-day soft delete, 7-year archive per policy (compliance to confirm).
- **Correlation keys (in priority order):** `employeeId` (HRIS) → `mail` (primary SMTP normalized) → `onPremisesImmutableId` → `userPrincipalName` → `displayName + hireDate` (advisory only, never auto-merge). See §7.
- **Namespaces:** `harborspecialty.com` and other Harbor vanity domains added to Tenant N as accepted domains and verified. UPN policy: Northstar keeps `@northstar.com`; migrated Harbor users retain existing UPN suffix. All Harbor SMTP addresses retained as proxy addresses.
- **Conditional Access baseline:** device compliance + phishing-resistant MFA for all users; risk-based sign-in policy; location-aware controls for US/UK/DE; separate policy set for kiosks (shared device sign-in, short session, no persistent tokens); separate policy for contractors (MAM-WE, block non-compliant clients).
- **Passwordless:** FIDO2 security keys for admins; Windows Hello for Business + Microsoft Authenticator (device-bound passkeys) for employees; extend to Harbor population as part of wave prerequisites.
- **PIM:** All privileged roles time-bound, approval-required, MFA-on-activation. Zero standing GA except two monitored break-glass. Access reviews quarterly.
- **Emergency access:** Two FIDO2-protected, monitored break-glass accounts, stored in split-knowledge safe. Alert on any sign-in.
- **Guest identities:** B2B collaboration for external partners only; cross-tenant access settings restricted to trusted tenants; entitlement management for guest lifecycle.
- **Logging:** Sign-in and audit logs to Log Analytics, 2-year hot, 7-year archive per compliance direction; Sentinel analytics for privileged actions and impossible travel.
- **Decommissioning boundaries:** Tenant H set to read-only 30 days after final wave; retained 90 days (Control 5); formal tenant closure at Day 120; retention exports handed to compliance before closure.

### Temporary coexistence controls (removed post-cutover)
- Cross-tenant sync (Tenant H → Tenant N) creating B2B member objects.
- Mail coexistence via shared namespace configuration and centralized mail transport (Tenant H accepted as internal relay for `harborspecialty.com` during transition).
- Dual SSO for the 3 object-ID-dependent apps (see §6) with identity bridge.
- Retained Harbor AD Connect until final Harbor user is cut over.
- Cross-tenant access settings allowing seamless Teams meetings and file access.
- Temporary CA policy allowing SMS MFA for non-privileged Harbor users during first 6 weeks after their cutover, with hard deadline for uplift.

---

## 4. 28-Week Phase Plan

Assumes program start = Week 0. Freezes: W9–W11 (quarter-end), W18–W21 (catastrophe response).

| Phase | Weeks | Entry criteria | Exit criteria | Key deliverables | Accountable | Pressure points |
|---|---|---|---|---|---|---|
| **P0 Mobilize** | 0–2 | Charter signed | RACI, vendor SOW draft, discovery kickoff | Program charter, RAID v1, comms plan | Program Manager | Vendor procurement lead time |
| **P1 Discover & Baseline** | 1–5 | P0 exit | Full inventory: users, groups, apps, mailboxes, devices, subscriptions, legal holds; correlation dataset v1 | Discovery report, correlation model, hold inventory | Identity Architects | Legal-hold discovery latency |
| **P2 Design & Vendor** | 3–6 | Discovery ≥80% complete | HLD/LLD signed; vendor onboarded; tooling deployed in lab | Target architecture, CA design, migration tooling PoC | Identity Architects + Vendor | Vendor onboarding 4-week runway |
| **P3 Foundation Build** | 5–8 | Design signed | CTS operational; accepted domains verified; PIM extended; break-glass tested; correlation ≥99.5% | Foundation runbook, CA policies deployed in report-only | Identity Engineers | Domain verification, DNS ownership |
| **⛔ Freeze 1** | 9–11 | — | No production identity changes | Regression tests, documentation, pilot prep | PM | Team morale; use for training |
| **P4 Pilot Wave** | 12–14 | Correlation ≥99.95% (Control 1); privileged uplift complete (Control 2) | Pilot cohort migrated with zero Sev-1; app owner sign-offs (Control 3) | Pilot report, tuned runbook | Identity Architects | Cohort composition |
| **P5 Privileged & Admin Migration** | 13–15 | Pilot pass; FIDO2 issuance complete for admins | All Harbor admins in Tenant N with PIM; standing GA = 0; shared emergency accounts retired | Admin migration report, PIM baseline | CISO delegate | Executive scheduling |
| **P6 Wave 1 (Corporate Non-Customer-Facing)** | 15–17 | P5 exit; legal-hold validation for cohort (Control 4) | 800 users migrated; stabilization thresholds met | Wave 1 closure report | Migration Lead | Freeze 2 approaching |
| **🟡 Restricted Window** | 18–21 | Catastrophe response | Only low-risk work: docs, contractor onboarding to B2B, app owner testing, non-production Azure subscription moves | Application readiness backlog burn-down | Migration Lead | Claims platform sacred; no customer-facing waves |
| **P7 Wave 2 (Customer-Facing incl. Claims)** | 22–24 | P6 closure; claims platform coexistence test passed | 2,200 users; 15-minute SLA maintained; kiosk cohort migrated | Wave 2 report | Migration Lead + Business Unit reps | Telephony dependency |
| **P8 Wave 3 (Remaining Employees + Contractors)** | 24–26 | P7 stabilization ≥72h clean | Remaining 1,600 employees + 900 contractors; BYOD onboarded | Wave 3 report | Migration Lead | Contractor identity provisioning |
| **P9 Azure + Residual Apps** | 25–27 | Users migrated | 12 subscriptions transferred; residual app cleanup; source tenant read-only | Subscription transfer records | Cloud lead (identity architect) | Billing agreements |
| **P10 Stabilize & Handover** | 27–28 | All waves complete | Ops runbook handover; scorecard closed; source tenant enters 90-day retention | Program closure report | PM | Documentation completeness |

### Pilot cohort composition (155 users, Week 12)
Chosen to exercise every migration pattern in a controlled sample:
- 25 IT staff (self-support, tolerant of issues)
- 20 knowledge workers, US (baseline)
- 20 knowledge workers, UK (regional CA)
- 15 knowledge workers, DE (data-residency, works-council touch)
- 15 claims adjusters, non-customer-facing shift (kiosk exposure)
- 10 users of one object-ID-dependent app (Tier-A)
- 10 users of an SCIM app + a SAML app
- 10 users with dual identity in both tenants (§7)
- 10 users with duplicate proxy addresses
- 10 shared mailbox delegates
- 5 executives (visibility)
- 5 contractors (B2B pattern)

**Representativeness rationale:** every collision class, every device class except macOS-heavy, every app class, both freeze-adjacent regions, and 3 language/regulatory contexts are exercised.

### Staffing pressure points
- **Weeks 5–8:** identity engineers over-allocated (foundation + design refinement). Vendor absorbs 40%.
- **Weeks 12–17:** application engineers are the bottleneck (per-app cutover testing).
- **Weeks 22–24:** service desk peak; contract 8 additional Tier-1 agents for 4 weeks.
- **Weeks 25–27:** endpoint engineers on kiosk and macOS re-enrollment; consider deferring macOS beyond W28 if slipping.

---

## 5. Migration-Wave Design

### Cohort sequencing

| Wave | Weeks | Population | Size | Rationale |
|---|---|---|---|---|
| Pilot | 12–14 | Mixed representative | 155 | Prove runbook |
| Admin | 13–15 | Privileged & shared accounts | ~90 | Control 2; must precede business waves |
| W1 | 15–17 | Corporate back-office (finance, HR, legal, IT-remainder) | 800 | Low customer impact, high internal visibility |
| Restricted | 18–21 | Contractors onboarded to B2B only; no cutover | 900 (staged) | Freeze compliance |
| W2a | 22 | Claims adjusters non-kiosk | 900 | Business-critical, in-hours SLA |
| W2b | 23 | Claims adjusters on kiosks | 480 kiosk sessions | Kiosk pattern; scheduled per shift |
| W2c | 24 | Customer service + telephony users | 820 | Telephony coordination |
| W3a | 25 | Remaining employees | 1,000 | Long tail |
| W3b | 26 | Remaining employees + contractors | 600 + 900 | Cleanup |

### Device class treatment
- **Windows laptops (7,900 total; ~3,100 Harbor):** Autopilot re-enrollment using ESP; user data preserved via OneDrive Known Folder Move confirmed before cutover.
- **macOS (2,300 total; ~900 Harbor):** ADE re-enrollment; Platform SSO with new tenant; scheduled during user off-hours.
- **Managed mobile (3,100):** Company Portal re-enroll; MAM policy replacement.
- **Kiosks (480):** Shared-device mode; wave in one facility per night; local support present.
- **Contractor BYOD (760):** No MDM; B2B guest or member with MAM-WE and CA compliance grant.

### Go/no-go thresholds (each wave)

| Threshold | Value | Action if breached |
|---|---|---|
| Correlation accuracy for cohort | ≥99.95% | Delay wave |
| Legal-hold validation completeness | 100% | Delay wave |
| App owner sign-off for cohort's apps | 100% of Tier-A, ≥95% Tier-B | Delay wave |
| Privileged users on FIDO2 | 100% of cohort admins | Delay wave |
| Pilot Sev-1 defects open | 0 | Delay wave |
| Service desk staffing | ≥120% of forecast | Delay wave |
| Freeze exception approved (if applicable) | Written approval | Cancel wave |

### Stabilization protocol per wave

| Interval | Actions | Owner | Exit |
|---|---|---|---|
| T+0 to T+24h | Hyper-care bridge open; sign-in success rate monitored per 15 min; mail flow test batches every hour; top-10 error triage | Migration Lead + Service Desk | <2% ticket-per-user rate; no Sev-1 |
| T+24h to T+72h | Daily reconciliation report (accounts, licenses, group memberships, mailbox status); app usage anomaly check | Identity Engineer on-call | All reconciliation deltas explained |
| T+72h to T+7d | Delegated permissions verification; retention label validation; guest access smoke test; wave retrospective | Migration Lead | Retrospective actions logged; wave closure signed by CISO delegate |

Rollback decision points defined per wave in §9.

---

## 6. Application Disposition

### Classification

| Class | Count | Pattern | Effort |
|---|---|---|---|
| SAML/OIDC re-federation, low complexity | ~50 of 73 | Add Tenant N as second IdP → cutover metadata → decommission Tenant H app | S |
| SAML/OIDC with claim/attribute customization | ~20 of 73 | As above + claims re-mapping; regression test | M |
| SAML/OIDC with hard IdP entity ID lock | ~3 of 73 | Vendor ticket to change entity ID; may need re-registration; schedule per-app cutover | L |
| SCIM provisioning | 18 | Re-point SCIM endpoint to Tenant N; rotate tokens; scope-test; run parallel provisioning for 48h; then disable Tenant H | M |
| LDAP via on-prem AD | 34 | **Defer.** Retain Harbor AD forest (or replace with resource forest) trusted from Northstar. Re-platform on 6–9 month follow-on. | Out of scope |
| Object-ID dependent, vendor-supported remap | 4 of 7 | Export mapping (old objectId → new objectId), run vendor API remap during cutover window, verify | M |
| Object-ID dependent, **no vendor remap** | 3 of 7 | See below | H |
| Azure subscriptions | 12 | Transfer to Tenant N EA/MCA; re-do RBAC from group entitlements (do not migrate identities as-is) | M |

### Three apps without bulk remap — practical strategy

Assumption: no undocumented vendor capability exists. Options in decreasing order of preference:

**Option A — Identity Bridge (recommended for Tier-A):**
- Keep app federated to Tenant H IdP during coexistence; add Tenant N users as B2B guests to Tenant H for authentication only.
- User experience: single sign-on via Tenant N credentials proxied through B2B; app sees original Tenant H `objectId`.
- Trade-off: preserves object-ID; delays true consolidation of that app to a follow-on project; requires Tenant H to remain live for authentication (adjust Control 5 to authentication-only, not full read-only).
- Escalation trigger: if Tenant H cannot remain live for authentication (e.g., licensing cost, security posture), move to Option B.

**Option B — Manual per-user re-linking with scheduled downtime:**
- Coordinate app owner + vendor for maintenance window; export user list; on cutover, admin creates a new record in the app for each Tenant N identity; historical data re-associated via a mapping table maintained in the app database (requires vendor read/write DB agreement).
- Trade-off: downtime per user or per batch; risk of lost historical linkage; vendor may not warrant DB manipulation.
- Escalation trigger: any vendor refusal → Option C.

**Option C — Replace the application:**
- Time-boxed evaluation (30 days) of alternatives; if replacement identified, put app in the 6–9 month follow-on backlog and continue coexistence via Option A.

**Recommendation:** Apply Option A for all three during the 28-week window; commit to a Q+2 project to eliminate the bridge. This preserves the schedule and Control 1 without pretending vendor capability exists.

---

## 7. Identity Correlation & Collision Resolution

### Matching hierarchy

| Rank | Key | Confidence | Action |
|---|---|---|---|
| 1 | `employeeId` present in both tenants and equal | High | Auto-correlate |
| 2 | Primary SMTP normalized match AND HR record match | High | Auto-correlate |
| 3 | `onPremisesImmutableId` match | High | Auto-correlate |
| 4 | UPN local part + hire date + manager match | Medium | Auto-correlate with sampling audit |
| 5 | `displayName + department + country` match | Low | Route to exception queue; human decision |
| 6 | Fuzzy match (Jaro-Winkler ≥0.92 on name + shared attributes) | Advisory only | Never auto-merge; suggest to reviewer |

### Category treatment

| Category | Volume | Treatment |
|---|---|---|
| Dual identity (person in both) | 46 | Person-by-person interview via HR partner; retain Tenant N identity as primary, migrate Tenant H mailbox contents, add Harbor SMTP as alias, transfer group memberships explicitly |
| Duplicate proxy addresses | 87 | Compare owners; if same person → merge; if different → keep Harbor SMTP on Harbor account only, rename Tenant N proxy per naming standard, notify affected users 14 days before cutover |
| Duplicate display names | 640 | Automatic disambiguation using department suffix; no name change required; documented in comms |
| Guest collisions | 1,200 | Delete stale guests (>90 days inactive) after owner sign-off; migrate active guests via CTS; re-invite where guest was a Harbor employee |
| Contractors | 900 | Fresh B2B member creation in Tenant N with sponsor attribute; no correlation to Tenant N employees |
| Shared mailboxes | Discovery pending | Recreate in Tenant N; migrate content; reapply delegate permissions from export; do not convert to user mailboxes |
| Recent name changes | Discovery pending | Match on `employeeId` + previous UPN captured from Entra audit logs |

### Control totals
Before every wave, publish reconciliation showing:
`in-scope users` = `auto-correlated` + `manually resolved` + `explicit exceptions with owner`
Any residual = wave blocked.

### Exception record format

```
ExceptionID: HAR-EX-0007
Type: DualIdentity | DuplicateProxy | GuestCollision | Other
SourceObjectId (Tenant H): <guid>
TargetObjectId (Tenant N): <guid or NEW>
MatchConfidence: High | Medium | Low | None
ProposedResolution: Merge | KeepBoth | RenameProxy | ManualCreate
Owner: <role, e.g., HR Business Partner - EMEA>
DecisionDeadline: <YYYY-MM-DD>
ApprovedBy: <role>
EvidenceLinks: [ticket, HR record, email]
Notes:
```

---

## 8. Security & Compliance Control Matrix

| Objective | Implementation | Evidence | Owner | Validation point | Failure response |
|---|---|---|---|---|---|
| Least-privileged admin | PIM eligible-only; 0 standing GA except 2 break-glass | PIM export weekly | CISO delegate | Weekly and pre-wave | Suspend privileged access; incident |
| Phishing-resistant MFA for admins | FIDO2 required for privileged; CA policy | CA policy export; sign-in logs | Identity Architect | Pre-P5 gate | Block admin migration |
| Separation of duties | Migration engineers cannot approve their own change; PIM approver ≠ requester; wave sign-off by CISO delegate | Change tickets; PIM approvals | PM | Every change | Change reversed |
| Audit log retention | Log Analytics 2y hot, 7y archive | Retention config screenshot | Security Analyst | Weekly | Escalate to CISO |
| Legal holds validated pre-migration | Compliance queries mailbox and site holds; sign-off per user cohort | Hold inventory report + sign-off | Compliance Lead + Counsel | Pre-wave gate (Control 4) | Wave halted |
| Data residency | Multi-geo EXO and SharePoint honored; user residency attribute drives placement | Multi-geo config; sample validation | Compliance Lead | Pre-cutover per region | User held back |
| Consent (DE works council, contractor privacy) | Privacy officer approval on data flow; DPIA update | DPIA doc; approval email | Compliance Lead + Privacy Officer | Pre-P4 | Region scope delayed |
| Dormant accounts | Disable accounts inactive >90d before migration | Inactivity report | Identity Engineer | Pre-wave | Excluded from wave |
| Emergency access | 2 break-glass FIDO2, split-knowledge, alert on sign-in | Vault log; alert config | CISO delegate | Monthly test | Rotate immediately |
| Source-tenant 90-day period | Tenant H set read-only; only 2 named admins retain access via PIM; all sign-ins alerted | Tenant H CA config; sign-in logs | Identity Architect | Continuous | Freeze access; investigate |
| Change control | All migration changes via CAB-lite with pre-approved templates | CAB records | PM | Weekly | Rollback |
| SMS MFA sunset | Temporary 6-week window post-cutover per user; forced enrollment to Authenticator/FIDO2 | Auth method report | Identity Engineer | 6 weeks post-wave | Account MFA reset |

**Matters requiring counsel or privacy-officer approval (not decided here):**
- Custodianship of legal holds when mailboxes move between tenants.
- Cross-border processing of German employee data during coexistence.
- Works-council consultation requirements in Germany.
- Contractor personal-device data-processing scope for B2B members vs. guests.
- Retention obligations for the Tenant H tenant itself post-decommission.
- Records-retention treatment of correlation exception decisions.

---

## 9. Cutover Runbook — Representative 500-User Wave

Times are in the wave's primary business time zone. Roles are named by function, not person.

### T-7 days
- Comms: Manager-to-user reminder issued (Service Desk Comms Lead).
- Correlation report re-run for cohort; must be 100% resolved (Identity Engineer).
- Legal-hold validation confirmed for cohort mailboxes (Compliance Lead).
- App owner sign-offs collected for all Tier-A and Tier-B apps in cohort scope (App Engineering Lead).
- Device readiness report: OneDrive KFM confirmed for each Windows user (Endpoint Lead).

### T-3 days
- Freeze secondary changes (Change Manager).
- Hypercare bridge scheduled; on-call rota confirmed (Migration Lead).
- Reversal test: verify last known-good backup of source objects (Identity Engineer).
- Executive brief sent (PM).

### T-1 day
- Pre-flight checklist reviewed live in go/no-go call at 14:00 local (Migration Lead chairs).
- **Latest safe decision point: T-1 day 18:00.** If aborted after this, communications sent within 1 hour.
- Break-glass account tested (CISO delegate).
- Service desk pre-brief and macro loading (Service Desk Manager).

### T-0 (cutover Friday 20:00 → Saturday 06:00)
- 20:00 Freeze source-tenant provisioning (Identity Engineer).
- 20:15 Disable password writeback on Harbor Connect (Identity Engineer).
- 20:30 Trigger CTS delta run; verify object counts (Identity Engineer).
- 21:00 Mailbox move requests submitted in batches of 50 (Messaging Engineer).
- 22:00 SCIM re-point for cohort apps (App Engineer).
- 23:00 SAML metadata cutover for cohort apps (App Engineer).
- 00:00 Intune re-enroll batch commences (Endpoint Engineer).
- 03:00 Smoke test suite executes (Test Lead).
- 04:00 Mail routing update for cohort's proxy addresses (Messaging Engineer).
- 05:00 Reconciliation control totals published (Identity Engineer).
- 05:30 Go/no-go for opening business (Migration Lead + CISO delegate + Business Unit rep).
- 06:00 Hypercare bridge remains open; business go-live.

**Rollback decision points (T-0):**
| Time | Trigger | Decision role | Recovery time | Post-cutover writes handling |
|---|---|---|---|---|
| 22:30 | Mailbox move failure rate >5% | Migration Lead | 2h to source state | None yet; no data loss |
| 00:30 | SSO failure for >2 apps | Migration Lead | 3h to source | None yet |
| 03:30 | Smoke test critical fail | Migration Lead + CISO delegate | 4h; users remain in Tenant H Monday | None yet |
| 05:30 | Reconciliation deltas unexplained >0.05% | Migration Lead + CISO delegate | 4h | None yet |
| **Post 06:00** | Sev-1 emerges after go-live | CIO delegate | Rollback becomes forward-fix; roll back only individual users; source-tenant writes captured via delta report and replayed manually | Written to Tenant N; source is read-only; per-user rollback recreates account and replays via export |

### T+24h
- Reconciliation report v2 (Identity Engineer).
- Ticket triage summary (Service Desk Manager).
- Sev-1/Sev-2 review at 10:00 (Migration Lead).

### T+72h
- Delegated permissions verification (Messaging Engineer).
- Group membership audit (Identity Engineer).
- App owner post-cutover confirmation (App Engineer).

### T+7d
- Wave retrospective (PM).
- Formal closure sign-off (CISO delegate).
- Lessons applied to next wave runbook (Migration Lead).

---

## 10. Test & Acceptance Plan

| Test class | Scope | Acceptance threshold |
|---|---|---|
| Positive functional | Sign-in, mail send/receive, Teams chat/meet/call, SPO access, Intune app deploy | 100% pass on smoke suite |
| Negative | Blocked-country sign-in, non-compliant device, expired MFA method | 100% blocked as designed |
| Performance | Sign-in latency, mailbox move throughput, provisioning latency | Sign-in p95 <3s; provisioning <15 min end-to-end |
| Security | Privileged access review, CA bypass attempt, token replay, break-glass alert | 0 successful bypass; 100% alert fire |
| Lifecycle (JML) | New hire, transfer, termination in migrated population | Provisioning ≤4h; termination ≤1h |
| Application | Per-app: authN, authZ, provisioning, deprovisioning, logging, rollback (Control 3) | 100% owner sign-off before wave |
| Device | Autopilot, ADE, mobile enroll, kiosk shared-mode | ≥98% success on first attempt |
| Coexistence | Cross-tenant Teams, calendar free/busy, mail flow both directions | 100% pass on 25-scenario matrix |
| DR | Simulated tenant-level auth outage; break-glass exercised | Access restored <30 min |
| Rollback | Per-wave dry run in lab; per-user rollback in production | Dry run: full source restore <4h; per-user: <1h |

Test evidence is filed against wave record; missing evidence = no wave sign-off.

---

## 11. Support & Communications Plan

### Audience matrix

| Audience | Message | Cadence | Owner |
|---|---|---|---|
| Executives | Scorecard, risk deltas, decisions needed | Weekly (60-min) | PM |
| Managers | Wave schedule for their teams, actions required | 4, 2, 1 weeks pre-wave | Business Unit reps |
| End users | What is changing, when, what to do | 14, 7, 1 day pre-wave; day-of; +2 days | Service Desk Comms Lead |
| Application owners | Test schedule, sign-off requests, cutover role | Continuous per app | App Engineering Lead |
| Service desk | Runbooks, macros, escalation | Weekly plus pre-wave brief | Service Desk Manager |
| SecOps | Detection tuning, wave anomalies expected | Pre-wave and daily during hypercare | Security Analyst |
| Regional privacy stakeholders (US/UK/DE) | Data-flow diagrams, DPIA updates, works-council notes | Milestone-based | Compliance Lead |
| Legal | Hold custodianship, tenant retention, contract items | Milestone-based | Compliance Lead |

### Service desk staffing assumption
- Baseline capacity handles 2,000 tickets/week.
- Wave surge assumption: 12% of migrated users create a ticket; 3% Sev-2+; peak day at T+1 business day.
- For W2 (2,200 users): +8 Tier-1 agents for 4 weeks, +2 Tier-2 for 6 weeks. German-language coverage required during DE cohort waves.

### Escalation path
Sev-3 → Tier-1 → Tier-2.
Sev-2 → Migration Lead within 30 min.
Sev-1 → Migration Lead + CISO delegate + PM within 15 min; CIO in loop within 30 min; hypercare bridge activated.

### Message timing (subject-line themes; drafts owned by Comms)
- T-14d: "What is changing and when"
- T-7d: "Your action items before Friday"
- T-1d: "Tonight's timeline and support numbers"
- T+0: "You are now on the new platform — first steps"
- T+2d: "Common questions and fixes"

---

## 12. RAID Register

R = Risk, A = Assumption, I = Issue, D = Dependency. Probability (P) and Impact (I) on 1–5.

| # | Type | Item | P | I | Leading indicator | Mitigation | Contingency | Owner | Decision deadline |
|---|---|---|---|---|---|---|---|---|---|
| 1 | R | 3 object-ID apps with no bulk remap | 5 | 5 | Vendor confirms no API | Option A identity bridge | Option B downtime; Option C replace | App Eng Lead | W3 |
| 2 | R | Privileged sprawl in Tenant H (39 GA) | 4 | 5 | GA audit >2 | Enforce PIM, retire shared accounts | Emergency access revocation | CISO delegate | W2 |
| 3 | R | Correlation misses 99.95% | 4 | 5 | Weekly correlation trend | HR partner desk; exception SLAs | Slip wave 2 weeks | Identity Architect | Rolling |
| 4 | R | Mail-flow outage at cutover | 3 | 5 | Failed dry runs | Rehearsals; smart-host; fallback MX | Revert mail routing (<30 min) | Messaging Eng | Pre-wave |
| 5 | R | Freeze exception denied for critical fix | 3 | 4 | CAB backlog | Pre-approved emergency template | Halt wave | PM | W8 |
| 6 | R | Vendor onboarding slippage | 3 | 4 | Contract redlines W3 | Parallel two-vendor evaluation | In-house scripts for CTS + reporting | PM | W2 |
| 7 | R | Kiosk shared-mode conflict with claims platform | 3 | 4 | Pilot kiosk test fail | Isolated pilot facility | Delay kiosk cohort to post-W28 | Endpoint Lead | W12 |
| 8 | R | DE works-council delay | 3 | 4 | Meeting minutes | Early engagement W2 | Descope DE from W2 wave | Compliance Lead | W6 |
| 9 | R | 15-min SLA breach during W2 | 2 | 5 | Pilot recovery time >15 min | Pre-staged rollback; per-user cutover for claims | Individual rollback | Migration Lead | W20 |
| 10 | R | Contractor B2B lifecycle unmanaged | 3 | 3 | No sponsor assigned in HR feed | Sponsor mandatory attribute; entitlement mgmt | Bulk disable at 30 days no sponsor | Identity Eng | W10 |
| 11 | R | Legal-hold gap discovered mid-wave | 2 | 5 | Sample audit fails | Pre-wave 100% inventory (Control 4) | Halt affected mailbox moves | Compliance Lead | Pre-wave |
| 12 | R | SharePoint permission drift post-migrate | 3 | 3 | Reconciliation deltas | Export/re-apply script; owner audit | Owner remediation SLA 5 days | App Eng Lead | Per-wave |
| 13 | R | Azure subscription RBAC re-mapped incorrectly | 3 | 4 | Missing group entitlements | Pre-build target group model; parallel run | Rollback subscription move (billing action) | Cloud Lead | Pre-P9 |
| 14 | R | Passwordless enrollment lag in Harbor | 4 | 3 | Enrollment <50% at W10 | Pop-up clinics; kit distribution | Extend temporary SMS window per user (not admin) | Identity Eng | W10 |
| 15 | R | SecOps drowned in benign anomalies during waves | 3 | 3 | Alert volume >2× baseline | Pre-wave detection tuning; expected-anomaly runbook | Suppress specific rules with sunset date | Security Analyst | Pre-wave |
| 16 | A | Harbor forests can remain during coexistence | — | — | AD forest health report | Confirm W1 | Retire on follow-on track | Identity Architect | W3 |
| 17 | D | Vendor SOW signed by W4 | — | — | Contract status | Legal fast-track | In-house fallback plan | PM | W3 |
| 18 | I | Six shared emergency accounts in Tenant H (Control 2) | 5 | 5 | Existing | Retire before W5; create individual PIM-eligible admins | Rotate credentials immediately if retention required | CISO delegate | W3 |

---

## 13. Budget Allocation

Incremental envelope: **$180,000**.

### Base case
| Line | Cost | Rationale |
|---|---|---|
| Specialist migration vendor (tooling + advisory) | $95,000 | CTS accelerator, mailbox/SPO movement, correlation reporting, 12 weeks |
| FIDO2 security keys (admins + high-risk roles, ~600 keys @ $55) | $33,000 | Control 2; buffer for loss |
| Kiosk shared-mode licensing gap / device certs | $8,000 | 480 kiosks |
| Correlation & reconciliation SaaS (identity governance add-on trial + entitlement mgmt tuning) | $12,000 | Exception queue tooling |
| Third-party mailbox coexistence add-on (fallback licensing) | $10,000 | Only if native path insufficient in pilot |
| Contingency reserve | $22,000 | 12% |
| **Total** | **$180,000** | |

### Contingency case (if vendor slips or 3 object-ID apps require Option B)
- Move $22K reserve + $10K coexistence line into: additional application-engineering contractor hours ($20K), extended service-desk augmentation ($12K).
- Trade-off: no third-party mail fallback; must rely on native path performing.

### Assumptions
- Existing Microsoft licenses cover Entra ID P2, Intune, Defender, Sentinel.
- German-language service desk capacity is procurable within SLA.
- Vendor pricing is within typical SOW ranges; a competitive selection is possible in 3 weeks.

### Where staffing is the binding constraint (not money)
- **Application engineering (3 FTE)** during W12–W17: even with unlimited budget, ramp-up time makes contract hires marginal. Recommend cross-training two identity engineers on SAML/SCIM basics in P0.
- **Endpoint engineering (3 FTE)** during W22–W26: kiosk + macOS + Windows re-enrollment concurrent; consider deferring macOS cohort into follow-on.
- **Compliance lead (1 FTE)** across all pre-wave gates: single point of failure for legal-hold sign-off. Nominate a backup with formal delegation from Counsel.

---

## 14. Program Scorecard

| # | Metric | Formula | Source | Cadence | Target | Warning | Executive interpretation |
|---|---|---|---|---|---|---|---|
| 1 | Correlation accuracy | Auto+resolved / total in scope | Correlation tool | Weekly | ≥99.95% | <99.90% | Wave gate |
| 2 | Privileged accounts on FIDO2 | FIDO2-enrolled admins / total admins | Entra auth methods | Weekly | 100% | <100% | Control 2 breach |
| 3 | Standing GA count | Direct query | Entra PIM | Weekly | 2 (break-glass) | ≥3 | Immediate review |
| 4 | Wave completion vs plan | Users migrated / plan | Migration tool | Weekly | 100% cumulative | −5% | Schedule risk |
| 5 | Sev-1 incidents per wave | Count | ITSM | Per wave | 0 | ≥1 | Rollback consideration |
| 6 | Business-hour outage minutes (customer-facing) | Cumulative minutes | ITSM | Per wave | 0 | ≥15 | SLA breach |
| 7 | Ticket rate per 100 migrated users | Tickets / users × 100 | ITSM | Daily during hypercare | <12 | >18 | Support capacity |
| 8 | Legal-hold validation coverage | Validated / migrated cohort | Compliance tool | Per wave | 100% | <100% | Wave halted |
| 9 | App owner sign-off coverage | Signed / cohort apps | Program tracker | Per wave | 100% Tier-A; ≥95% Tier-B | Below | Wave halted |
| 10 | Passwordless adoption (migrated Harbor) | Passwordless users / migrated | Auth methods | Weekly | ≥70% by W+6 post-wave | <50% | Adoption push |
| 11 | Source-tenant sign-ins during 90d retention | Sign-in count from named admins only | Audit logs | Weekly | Only 2 named admins | Any other | Investigate |
| 12 | Budget consumed vs plan | Spent / planned | Finance | Bi-weekly | ≤105% | >110% | Rebaseline |

---

## 15. Steering Committee One-Page Checklist

**Approvals required at next steering meeting**
- [ ] D1 — Endorse scoped migration pattern (LDAP deferred to follow-on)
- [ ] D2 — Approve vendor selection and $180K budget envelope
- [ ] D3 — Mandate retirement of Harbor shared emergency accounts and FIDO2 for admins by W5
- [ ] D4 — Confirm Tenant N as destination; approve Harbor domain addition
- [ ] D5 — Counsel and Privacy Officer confirmations for cross-border data flow and legal-hold custodianship
- [ ] D6 — Approve contractor identity model and BYOD policy
- [ ] D7 — Name single freeze-exception approver and go/no-go authority
- [ ] Approve wave schedule respecting quarter-end (W9–W11) and catastrophe (W18–W21) freezes
- [ ] Approve identity bridge (Option A) for three non-remappable applications with Q+2 exit plan
- [ ] Approve 90-day Tenant H retention posture with 2 named admins

**Unresolved decisions to schedule**
- Follow-on program budget and timeline for 34 LDAP apps
- Follow-on decision to eliminate identity bridge on 3 object-ID apps
- Final DE works-council engagement scope and timing
- macOS deferral trigger criteria if W25–W26 capacity is exceeded
- Retention treatment of correlation exception records (records-retention policy interpretation → Counsel)
- Post-program identity governance operating model (steady-state RACI)

**Key dates to note**
- W2: D1–D7 decisions due
- W4: Vendor onboarding complete
- W8: Foundation exit gate (pre-freeze)
- W12: Correlation gate (Control 1) and pilot launch
- W15: Privileged migration complete (Control 2)
- W22–W24: Customer-facing wave (15-minute SLA in force)
- W27: Source tenant enters 90-day retention
- W+90 days after final wave: Tenant H closure decision

---

*End of plan. Prepared for decision, not for record. Update after each steering meeting; re-baseline scorecard at end of each phase.*
Finish reason: stop
Try in Playground
import { createClient } from '@runware/sdk'

const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()

const [result] = await client.run({
  model: 'anthropic:claude@opus-4.7',
  settings: {
    thinkingLevel: 'high'
  },
  messages: [
    {
      role: 'user',
      content: 'Act as the principal identity architect for a regulated insurance group. Produce the execution plan that the CIO, CISO, service owners, and migration team will use to consolidate two Microsoft Entra ID tenants after an acquisition.\n\nScenario\n- Northstar Mutual has 11,800 employees in tenant N and acquired Harbor Specialty, which has 4,600 employees and 900 contractors in tenant H.\n- Tenant N will be the destination tenant. The program must finish within 28 weeks.\n- Both companies use Microsoft 365, Exchange Online, Teams, SharePoint, Intune, and Entra ID. Tenant H also has 73 enterprise applications using SAML or OIDC, 18 applications using SCIM provisioning, 34 legacy applications using LDAP through on-premises Active Directory, and 12 Azure subscriptions.\n- Tenant N uses cloud-native authentication with passwordless sign-in for 68% of employees. Tenant H synchronizes identities from two on-premises Active Directory forests and relies heavily on SMS MFA.\n- Harbor employees must retain their existing email addresses as aliases. Customer-facing staff cannot lose access to email, Teams, the claims platform, or telephony for more than 15 minutes during business hours.\n- Seven applications identify users by immutable object ID rather than email address. Three of those applications have no vendor-supported bulk remapping function.\n- There are approximately 1,200 guest-account collisions, 640 duplicate display names, 87 duplicate proxy addresses, and 46 people who already hold identities in both tenants.\n- Privileged roles are currently inconsistent. Tenant H has 39 Global Administrators, including 11 standing assignments and 6 shared emergency accounts. Tenant N uses Privileged Identity Management and maintains two monitored break-glass accounts.\n- Devices include 7,900 Windows laptops, 2,300 macOS devices, 3,100 managed mobile devices, 480 shared clinical-style kiosks used by claims adjusters, and 760 personally owned contractor devices.\n- The company operates in the United States, United Kingdom, and Germany. The plan must respect least privilege, auditable approvals, data-residency obligations, legal holds, records-retention policies, and separation of duties. Do not provide legal conclusions; identify matters requiring counsel or privacy-officer approval.\n- Quarter-end runs from weeks 9 through 11, and no production identity changes are permitted during that freeze. The claims catastrophe-response period runs from weeks 18 through 21; only low-risk migration work may occur then.\n- Available team: one program manager, two identity architects, four identity engineers, three endpoint engineers, two messaging engineers, three application engineers, two security analysts, one compliance lead, and part-time representatives from the service desk, HR, legal, and each business unit.\n- Approved tooling budget beyond existing licenses is $180,000. A specialist migration vendor may be engaged, but vendor selection and onboarding require four weeks.\n\nNon-negotiable controls\n1. No bulk migration wave may begin until identity correlation reaches at least 99.95% accuracy and every unresolved collision has a named owner.\n2. Every privileged identity must use phishing-resistant MFA before it enters the destination tenant.\n3. Application owners must sign off on authentication, authorization, provisioning, deprovisioning, logging, and rollback tests.\n4. Legal holds and retention labels must be validated before mailbox or SharePoint migration.\n5. The source tenant must remain recoverable and read-only for at least 90 days after the final user wave.\n6. Rollback must be defined in operational terms, including the latest safe decision point, responsible role, expected recovery time, and treatment of writes made after cutover.\n\nCreate a decision-ready execution plan in Markdown. It must be specific enough for teams to schedule and operate from, not a generic identity-migration overview. Include:\n\n1. Executive recommendation: the proposed migration pattern, why it is preferable to plausible alternatives, the five largest risks, and the decisions executives must make within the next two weeks.\n2. Current-state problem decomposition: identity authority, authentication, authorization, lifecycle management, applications, messaging and collaboration, endpoints, Azure resources, privileged access, compliance, and support readiness.\n3. Target-state architecture: authoritative sources, joiner/mover/leaver flow, account-correlation keys, domain and namespace handling, cross-tenant coexistence, Conditional Access, passwordless authentication, PIM, emergency access, guest identities, logging, and decommissioning boundaries. Clearly distinguish temporary coexistence controls from the permanent design.\n4. A 28-week phase plan that accounts for both change freezes. Show phases, entry and exit criteria, dependencies, accountable roles, staffing pressure points, and tangible deliverables. Include pilot cohorts and explain why they are representative.\n5. Migration-wave design: propose cohort sizes and sequencing for employees, contractors, privileged users, shared accounts, kiosks, and each major device class. Define go/no-go thresholds and a 24-hour, 72-hour, and 7-day stabilization process for every wave.\n6. Application disposition: classify the 73 federated applications, 18 SCIM integrations, 34 LDAP applications, and seven object-ID-dependent applications into migration patterns. Provide a practical strategy for the three applications that cannot bulk-remap users, including trade-offs and escalation triggers. Do not assume an undocumented vendor capability.\n7. Identity-correlation and collision-resolution procedure: matching hierarchy, confidence levels, treatment of dual identities, duplicate aliases, guests, contractors, shared mailboxes, and people with name changes. Include a control-total approach and a sample exception-record format.\n8. Security and compliance control matrix with control objective, implementation, evidence, owner, validation point, and failure response. Cover privileged access, phishing-resistant MFA, separation of duties, audit retention, legal holds, data residency, consent, dormant accounts, emergency access, and source-tenant access during the 90-day recovery period.\n9. Operational cutover runbook for one representative 500-user wave. Give a time-ordered schedule from T-7 days through T+7 days, with checkpoints, communications, monitoring, reconciliation, rollback decision points, and named roles rather than individual names.\n10. Test and acceptance plan covering positive, negative, performance, security, lifecycle, application, device, coexistence, disaster-recovery, and rollback testing. Include measurable acceptance thresholds.\n11. Support and communications plan for executives, managers, end users, application owners, the service desk, security operations, and regional privacy stakeholders. Include service-desk staffing assumptions, escalation paths, and message timing, but do not draft lengthy email copy.\n12. RAID register containing at least 15 substantive items, each with probability, impact, leading indicator, mitigation, contingency, owner, and decision deadline.\n13. Budget allocation for the $180,000 incremental budget. Present a base case and a contingency case, state assumptions, and flag where staffing capacity rather than tooling is the binding constraint.\n14. Program scorecard with no more than 12 metrics. Define formula, data source, reporting cadence, target, warning threshold, and executive interpretation for each metric.\n15. A final one-page steering-committee checklist of approvals and unresolved decisions.\n\nUse concise prose, tables where they improve usability, and explicit recommendations. Identify assumptions and unknowns rather than silently inventing facts. Resolve tensions among schedule, security, reversibility, staffing, and user disruption. If the 28-week deadline is infeasible under the stated controls, say so clearly and present the smallest defensible scope or resource adjustment that would make it achievable.'
    }
  ]
})
import asyncio
import os

from runware import Runware


async def main():
    async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
        results = await client.run({
            "model": "anthropic:claude@opus-4.7",
            "settings": {
                "thinkingLevel": "high"
            },
            "messages": [
                {
                    "role": "user",
                    "content": "Act as the principal identity architect for a regulated insurance group. Produce the execution plan that the CIO, CISO, service owners, and migration team will use to consolidate two Microsoft Entra ID tenants after an acquisition.\n\nScenario\n- Northstar Mutual has 11,800 employees in tenant N and acquired Harbor Specialty, which has 4,600 employees and 900 contractors in tenant H.\n- Tenant N will be the destination tenant. The program must finish within 28 weeks.\n- Both companies use Microsoft 365, Exchange Online, Teams, SharePoint, Intune, and Entra ID. Tenant H also has 73 enterprise applications using SAML or OIDC, 18 applications using SCIM provisioning, 34 legacy applications using LDAP through on-premises Active Directory, and 12 Azure subscriptions.\n- Tenant N uses cloud-native authentication with passwordless sign-in for 68% of employees. Tenant H synchronizes identities from two on-premises Active Directory forests and relies heavily on SMS MFA.\n- Harbor employees must retain their existing email addresses as aliases. Customer-facing staff cannot lose access to email, Teams, the claims platform, or telephony for more than 15 minutes during business hours.\n- Seven applications identify users by immutable object ID rather than email address. Three of those applications have no vendor-supported bulk remapping function.\n- There are approximately 1,200 guest-account collisions, 640 duplicate display names, 87 duplicate proxy addresses, and 46 people who already hold identities in both tenants.\n- Privileged roles are currently inconsistent. Tenant H has 39 Global Administrators, including 11 standing assignments and 6 shared emergency accounts. Tenant N uses Privileged Identity Management and maintains two monitored break-glass accounts.\n- Devices include 7,900 Windows laptops, 2,300 macOS devices, 3,100 managed mobile devices, 480 shared clinical-style kiosks used by claims adjusters, and 760 personally owned contractor devices.\n- The company operates in the United States, United Kingdom, and Germany. The plan must respect least privilege, auditable approvals, data-residency obligations, legal holds, records-retention policies, and separation of duties. Do not provide legal conclusions; identify matters requiring counsel or privacy-officer approval.\n- Quarter-end runs from weeks 9 through 11, and no production identity changes are permitted during that freeze. The claims catastrophe-response period runs from weeks 18 through 21; only low-risk migration work may occur then.\n- Available team: one program manager, two identity architects, four identity engineers, three endpoint engineers, two messaging engineers, three application engineers, two security analysts, one compliance lead, and part-time representatives from the service desk, HR, legal, and each business unit.\n- Approved tooling budget beyond existing licenses is $180,000. A specialist migration vendor may be engaged, but vendor selection and onboarding require four weeks.\n\nNon-negotiable controls\n1. No bulk migration wave may begin until identity correlation reaches at least 99.95% accuracy and every unresolved collision has a named owner.\n2. Every privileged identity must use phishing-resistant MFA before it enters the destination tenant.\n3. Application owners must sign off on authentication, authorization, provisioning, deprovisioning, logging, and rollback tests.\n4. Legal holds and retention labels must be validated before mailbox or SharePoint migration.\n5. The source tenant must remain recoverable and read-only for at least 90 days after the final user wave.\n6. Rollback must be defined in operational terms, including the latest safe decision point, responsible role, expected recovery time, and treatment of writes made after cutover.\n\nCreate a decision-ready execution plan in Markdown. It must be specific enough for teams to schedule and operate from, not a generic identity-migration overview. Include:\n\n1. Executive recommendation: the proposed migration pattern, why it is preferable to plausible alternatives, the five largest risks, and the decisions executives must make within the next two weeks.\n2. Current-state problem decomposition: identity authority, authentication, authorization, lifecycle management, applications, messaging and collaboration, endpoints, Azure resources, privileged access, compliance, and support readiness.\n3. Target-state architecture: authoritative sources, joiner/mover/leaver flow, account-correlation keys, domain and namespace handling, cross-tenant coexistence, Conditional Access, passwordless authentication, PIM, emergency access, guest identities, logging, and decommissioning boundaries. Clearly distinguish temporary coexistence controls from the permanent design.\n4. A 28-week phase plan that accounts for both change freezes. Show phases, entry and exit criteria, dependencies, accountable roles, staffing pressure points, and tangible deliverables. Include pilot cohorts and explain why they are representative.\n5. Migration-wave design: propose cohort sizes and sequencing for employees, contractors, privileged users, shared accounts, kiosks, and each major device class. Define go/no-go thresholds and a 24-hour, 72-hour, and 7-day stabilization process for every wave.\n6. Application disposition: classify the 73 federated applications, 18 SCIM integrations, 34 LDAP applications, and seven object-ID-dependent applications into migration patterns. Provide a practical strategy for the three applications that cannot bulk-remap users, including trade-offs and escalation triggers. Do not assume an undocumented vendor capability.\n7. Identity-correlation and collision-resolution procedure: matching hierarchy, confidence levels, treatment of dual identities, duplicate aliases, guests, contractors, shared mailboxes, and people with name changes. Include a control-total approach and a sample exception-record format.\n8. Security and compliance control matrix with control objective, implementation, evidence, owner, validation point, and failure response. Cover privileged access, phishing-resistant MFA, separation of duties, audit retention, legal holds, data residency, consent, dormant accounts, emergency access, and source-tenant access during the 90-day recovery period.\n9. Operational cutover runbook for one representative 500-user wave. Give a time-ordered schedule from T-7 days through T+7 days, with checkpoints, communications, monitoring, reconciliation, rollback decision points, and named roles rather than individual names.\n10. Test and acceptance plan covering positive, negative, performance, security, lifecycle, application, device, coexistence, disaster-recovery, and rollback testing. Include measurable acceptance thresholds.\n11. Support and communications plan for executives, managers, end users, application owners, the service desk, security operations, and regional privacy stakeholders. Include service-desk staffing assumptions, escalation paths, and message timing, but do not draft lengthy email copy.\n12. RAID register containing at least 15 substantive items, each with probability, impact, leading indicator, mitigation, contingency, owner, and decision deadline.\n13. Budget allocation for the $180,000 incremental budget. Present a base case and a contingency case, state assumptions, and flag where staffing capacity rather than tooling is the binding constraint.\n14. Program scorecard with no more than 12 metrics. Define formula, data source, reporting cadence, target, warning threshold, and executive interpretation for each metric.\n15. A final one-page steering-committee checklist of approvals and unresolved decisions.\n\nUse concise prose, tables where they improve usability, and explicit recommendations. Identify assumptions and unknowns rather than silently inventing facts. Resolve tensions among schedule, security, reversibility, staffing, and user disruption. If the 28-week deadline is infeasible under the stated controls, say so clearly and present the smallest defensible scope or resource adjustment that would make it achievable."
                }
            ]
        })


asyncio.run(main())
curl https://api.runware.ai/v1 \
  -H "Authorization: Bearer $RUNWARE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '[
    {
      "taskType": "textInference",
      "taskUUID": "adbd5955-f339-4262-b369-1abc2788a08e",
      "model": "anthropic:claude@opus-4.7",
      "settings": {
        "thinkingLevel": "high"
      },
      "messages": [
        {
          "role": "user",
          "content": "Act as the principal identity architect for a regulated insurance group. Produce the execution plan that the CIO, CISO, service owners, and migration team will use to consolidate two Microsoft Entra ID tenants after an acquisition.\n\nScenario\n- Northstar Mutual has 11,800 employees in tenant N and acquired Harbor Specialty, which has 4,600 employees and 900 contractors in tenant H.\n- Tenant N will be the destination tenant. The program must finish within 28 weeks.\n- Both companies use Microsoft 365, Exchange Online, Teams, SharePoint, Intune, and Entra ID. Tenant H also has 73 enterprise applications using SAML or OIDC, 18 applications using SCIM provisioning, 34 legacy applications using LDAP through on-premises Active Directory, and 12 Azure subscriptions.\n- Tenant N uses cloud-native authentication with passwordless sign-in for 68% of employees. Tenant H synchronizes identities from two on-premises Active Directory forests and relies heavily on SMS MFA.\n- Harbor employees must retain their existing email addresses as aliases. Customer-facing staff cannot lose access to email, Teams, the claims platform, or telephony for more than 15 minutes during business hours.\n- Seven applications identify users by immutable object ID rather than email address. Three of those applications have no vendor-supported bulk remapping function.\n- There are approximately 1,200 guest-account collisions, 640 duplicate display names, 87 duplicate proxy addresses, and 46 people who already hold identities in both tenants.\n- Privileged roles are currently inconsistent. Tenant H has 39 Global Administrators, including 11 standing assignments and 6 shared emergency accounts. Tenant N uses Privileged Identity Management and maintains two monitored break-glass accounts.\n- Devices include 7,900 Windows laptops, 2,300 macOS devices, 3,100 managed mobile devices, 480 shared clinical-style kiosks used by claims adjusters, and 760 personally owned contractor devices.\n- The company operates in the United States, United Kingdom, and Germany. The plan must respect least privilege, auditable approvals, data-residency obligations, legal holds, records-retention policies, and separation of duties. Do not provide legal conclusions; identify matters requiring counsel or privacy-officer approval.\n- Quarter-end runs from weeks 9 through 11, and no production identity changes are permitted during that freeze. The claims catastrophe-response period runs from weeks 18 through 21; only low-risk migration work may occur then.\n- Available team: one program manager, two identity architects, four identity engineers, three endpoint engineers, two messaging engineers, three application engineers, two security analysts, one compliance lead, and part-time representatives from the service desk, HR, legal, and each business unit.\n- Approved tooling budget beyond existing licenses is $180,000. A specialist migration vendor may be engaged, but vendor selection and onboarding require four weeks.\n\nNon-negotiable controls\n1. No bulk migration wave may begin until identity correlation reaches at least 99.95% accuracy and every unresolved collision has a named owner.\n2. Every privileged identity must use phishing-resistant MFA before it enters the destination tenant.\n3. Application owners must sign off on authentication, authorization, provisioning, deprovisioning, logging, and rollback tests.\n4. Legal holds and retention labels must be validated before mailbox or SharePoint migration.\n5. The source tenant must remain recoverable and read-only for at least 90 days after the final user wave.\n6. Rollback must be defined in operational terms, including the latest safe decision point, responsible role, expected recovery time, and treatment of writes made after cutover.\n\nCreate a decision-ready execution plan in Markdown. It must be specific enough for teams to schedule and operate from, not a generic identity-migration overview. Include:\n\n1. Executive recommendation: the proposed migration pattern, why it is preferable to plausible alternatives, the five largest risks, and the decisions executives must make within the next two weeks.\n2. Current-state problem decomposition: identity authority, authentication, authorization, lifecycle management, applications, messaging and collaboration, endpoints, Azure resources, privileged access, compliance, and support readiness.\n3. Target-state architecture: authoritative sources, joiner/mover/leaver flow, account-correlation keys, domain and namespace handling, cross-tenant coexistence, Conditional Access, passwordless authentication, PIM, emergency access, guest identities, logging, and decommissioning boundaries. Clearly distinguish temporary coexistence controls from the permanent design.\n4. A 28-week phase plan that accounts for both change freezes. Show phases, entry and exit criteria, dependencies, accountable roles, staffing pressure points, and tangible deliverables. Include pilot cohorts and explain why they are representative.\n5. Migration-wave design: propose cohort sizes and sequencing for employees, contractors, privileged users, shared accounts, kiosks, and each major device class. Define go/no-go thresholds and a 24-hour, 72-hour, and 7-day stabilization process for every wave.\n6. Application disposition: classify the 73 federated applications, 18 SCIM integrations, 34 LDAP applications, and seven object-ID-dependent applications into migration patterns. Provide a practical strategy for the three applications that cannot bulk-remap users, including trade-offs and escalation triggers. Do not assume an undocumented vendor capability.\n7. Identity-correlation and collision-resolution procedure: matching hierarchy, confidence levels, treatment of dual identities, duplicate aliases, guests, contractors, shared mailboxes, and people with name changes. Include a control-total approach and a sample exception-record format.\n8. Security and compliance control matrix with control objective, implementation, evidence, owner, validation point, and failure response. Cover privileged access, phishing-resistant MFA, separation of duties, audit retention, legal holds, data residency, consent, dormant accounts, emergency access, and source-tenant access during the 90-day recovery period.\n9. Operational cutover runbook for one representative 500-user wave. Give a time-ordered schedule from T-7 days through T+7 days, with checkpoints, communications, monitoring, reconciliation, rollback decision points, and named roles rather than individual names.\n10. Test and acceptance plan covering positive, negative, performance, security, lifecycle, application, device, coexistence, disaster-recovery, and rollback testing. Include measurable acceptance thresholds.\n11. Support and communications plan for executives, managers, end users, application owners, the service desk, security operations, and regional privacy stakeholders. Include service-desk staffing assumptions, escalation paths, and message timing, but do not draft lengthy email copy.\n12. RAID register containing at least 15 substantive items, each with probability, impact, leading indicator, mitigation, contingency, owner, and decision deadline.\n13. Budget allocation for the $180,000 incremental budget. Present a base case and a contingency case, state assumptions, and flag where staffing capacity rather than tooling is the binding constraint.\n14. Program scorecard with no more than 12 metrics. Define formula, data source, reporting cadence, target, warning threshold, and executive interpretation for each metric.\n15. A final one-page steering-committee checklist of approvals and unresolved decisions.\n\nUse concise prose, tables where they improve usability, and explicit recommendations. Identify assumptions and unknowns rather than silently inventing facts. Resolve tensions among schedule, security, reversibility, staffing, and user disruption. If the 28-week deadline is infeasible under the stated controls, say so clearly and present the smallest defensible scope or resource adjustment that would make it achievable."
        }
      ]
    }
  ]'
runware run anthropic:claude@opus-4.7 \
  settings.thinkingLevel=high \
  messages.0.role=user \
  messages.0.content="Act as the principal identity architect for a regulated insurance group. Produce the execution plan that the CIO, CISO, service owners, and migration team will use to consolidate two Microsoft Entra ID tenants after an acquisition.

Scenario
- Northstar Mutual has 11,800 employees in tenant N and acquired Harbor Specialty, which has 4,600 employees and 900 contractors in tenant H.
- Tenant N will be the destination tenant. The program must finish within 28 weeks.
- Both companies use Microsoft 365, Exchange Online, Teams, SharePoint, Intune, and Entra ID. Tenant H also has 73 enterprise applications using SAML or OIDC, 18 applications using SCIM provisioning, 34 legacy applications using LDAP through on-premises Active Directory, and 12 Azure subscriptions.
- Tenant N uses cloud-native authentication with passwordless sign-in for 68% of employees. Tenant H synchronizes identities from two on-premises Active Directory forests and relies heavily on SMS MFA.
- Harbor employees must retain their existing email addresses as aliases. Customer-facing staff cannot lose access to email, Teams, the claims platform, or telephony for more than 15 minutes during business hours.
- Seven applications identify users by immutable object ID rather than email address. Three of those applications have no vendor-supported bulk remapping function.
- There are approximately 1,200 guest-account collisions, 640 duplicate display names, 87 duplicate proxy addresses, and 46 people who already hold identities in both tenants.
- Privileged roles are currently inconsistent. Tenant H has 39 Global Administrators, including 11 standing assignments and 6 shared emergency accounts. Tenant N uses Privileged Identity Management and maintains two monitored break-glass accounts.
- Devices include 7,900 Windows laptops, 2,300 macOS devices, 3,100 managed mobile devices, 480 shared clinical-style kiosks used by claims adjusters, and 760 personally owned contractor devices.
- The company operates in the United States, United Kingdom, and Germany. The plan must respect least privilege, auditable approvals, data-residency obligations, legal holds, records-retention policies, and separation of duties. Do not provide legal conclusions; identify matters requiring counsel or privacy-officer approval.
- Quarter-end runs from weeks 9 through 11, and no production identity changes are permitted during that freeze. The claims catastrophe-response period runs from weeks 18 through 21; only low-risk migration work may occur then.
- Available team: one program manager, two identity architects, four identity engineers, three endpoint engineers, two messaging engineers, three application engineers, two security analysts, one compliance lead, and part-time representatives from the service desk, HR, legal, and each business unit.
- Approved tooling budget beyond existing licenses is \$180,000. A specialist migration vendor may be engaged, but vendor selection and onboarding require four weeks.

Non-negotiable controls
1. No bulk migration wave may begin until identity correlation reaches at least 99.95% accuracy and every unresolved collision has a named owner.
2. Every privileged identity must use phishing-resistant MFA before it enters the destination tenant.
3. Application owners must sign off on authentication, authorization, provisioning, deprovisioning, logging, and rollback tests.
4. Legal holds and retention labels must be validated before mailbox or SharePoint migration.
5. The source tenant must remain recoverable and read-only for at least 90 days after the final user wave.
6. Rollback must be defined in operational terms, including the latest safe decision point, responsible role, expected recovery time, and treatment of writes made after cutover.

Create a decision-ready execution plan in Markdown. It must be specific enough for teams to schedule and operate from, not a generic identity-migration overview. Include:

1. Executive recommendation: the proposed migration pattern, why it is preferable to plausible alternatives, the five largest risks, and the decisions executives must make within the next two weeks.
2. Current-state problem decomposition: identity authority, authentication, authorization, lifecycle management, applications, messaging and collaboration, endpoints, Azure resources, privileged access, compliance, and support readiness.
3. Target-state architecture: authoritative sources, joiner/mover/leaver flow, account-correlation keys, domain and namespace handling, cross-tenant coexistence, Conditional Access, passwordless authentication, PIM, emergency access, guest identities, logging, and decommissioning boundaries. Clearly distinguish temporary coexistence controls from the permanent design.
4. A 28-week phase plan that accounts for both change freezes. Show phases, entry and exit criteria, dependencies, accountable roles, staffing pressure points, and tangible deliverables. Include pilot cohorts and explain why they are representative.
5. Migration-wave design: propose cohort sizes and sequencing for employees, contractors, privileged users, shared accounts, kiosks, and each major device class. Define go/no-go thresholds and a 24-hour, 72-hour, and 7-day stabilization process for every wave.
6. Application disposition: classify the 73 federated applications, 18 SCIM integrations, 34 LDAP applications, and seven object-ID-dependent applications into migration patterns. Provide a practical strategy for the three applications that cannot bulk-remap users, including trade-offs and escalation triggers. Do not assume an undocumented vendor capability.
7. Identity-correlation and collision-resolution procedure: matching hierarchy, confidence levels, treatment of dual identities, duplicate aliases, guests, contractors, shared mailboxes, and people with name changes. Include a control-total approach and a sample exception-record format.
8. Security and compliance control matrix with control objective, implementation, evidence, owner, validation point, and failure response. Cover privileged access, phishing-resistant MFA, separation of duties, audit retention, legal holds, data residency, consent, dormant accounts, emergency access, and source-tenant access during the 90-day recovery period.
9. Operational cutover runbook for one representative 500-user wave. Give a time-ordered schedule from T-7 days through T+7 days, with checkpoints, communications, monitoring, reconciliation, rollback decision points, and named roles rather than individual names.
10. Test and acceptance plan covering positive, negative, performance, security, lifecycle, application, device, coexistence, disaster-recovery, and rollback testing. Include measurable acceptance thresholds.
11. Support and communications plan for executives, managers, end users, application owners, the service desk, security operations, and regional privacy stakeholders. Include service-desk staffing assumptions, escalation paths, and message timing, but do not draft lengthy email copy.
12. RAID register containing at least 15 substantive items, each with probability, impact, leading indicator, mitigation, contingency, owner, and decision deadline.
13. Budget allocation for the \$180,000 incremental budget. Present a base case and a contingency case, state assumptions, and flag where staffing capacity rather than tooling is the binding constraint.
14. Program scorecard with no more than 12 metrics. Define formula, data source, reporting cadence, target, warning threshold, and executive interpretation for each metric.
15. A final one-page steering-committee checklist of approvals and unresolved decisions.

Use concise prose, tables where they improve usability, and explicit recommendations. Identify assumptions and unknowns rather than silently inventing facts. Resolve tensions among schedule, security, reversibility, staffing, and user disruption. If the 28-week deadline is infeasible under the stated controls, say so clearly and present the smallest defensible scope or resource adjustment that would make it achievable."
{
  "taskType": "textInference",
  "taskUUID": "adbd5955-f339-4262-b369-1abc2788a08e",
  "model": "anthropic:claude@opus-4.7",
  "settings": {
    "thinkingLevel": "high"
  },
  "messages": [
    {
      "role": "user",
      "content": "Act as the principal identity architect for a regulated insurance group. Produce the execution plan that the CIO, CISO, service owners, and migration team will use to consolidate two Microsoft Entra ID tenants after an acquisition.\n\nScenario\n- Northstar Mutual has 11,800 employees in tenant N and acquired Harbor Specialty, which has 4,600 employees and 900 contractors in tenant H.\n- Tenant N will be the destination tenant. The program must finish within 28 weeks.\n- Both companies use Microsoft 365, Exchange Online, Teams, SharePoint, Intune, and Entra ID. Tenant H also has 73 enterprise applications using SAML or OIDC, 18 applications using SCIM provisioning, 34 legacy applications using LDAP through on-premises Active Directory, and 12 Azure subscriptions.\n- Tenant N uses cloud-native authentication with passwordless sign-in for 68% of employees. Tenant H synchronizes identities from two on-premises Active Directory forests and relies heavily on SMS MFA.\n- Harbor employees must retain their existing email addresses as aliases. Customer-facing staff cannot lose access to email, Teams, the claims platform, or telephony for more than 15 minutes during business hours.\n- Seven applications identify users by immutable object ID rather than email address. Three of those applications have no vendor-supported bulk remapping function.\n- There are approximately 1,200 guest-account collisions, 640 duplicate display names, 87 duplicate proxy addresses, and 46 people who already hold identities in both tenants.\n- Privileged roles are currently inconsistent. Tenant H has 39 Global Administrators, including 11 standing assignments and 6 shared emergency accounts. Tenant N uses Privileged Identity Management and maintains two monitored break-glass accounts.\n- Devices include 7,900 Windows laptops, 2,300 macOS devices, 3,100 managed mobile devices, 480 shared clinical-style kiosks used by claims adjusters, and 760 personally owned contractor devices.\n- The company operates in the United States, United Kingdom, and Germany. The plan must respect least privilege, auditable approvals, data-residency obligations, legal holds, records-retention policies, and separation of duties. Do not provide legal conclusions; identify matters requiring counsel or privacy-officer approval.\n- Quarter-end runs from weeks 9 through 11, and no production identity changes are permitted during that freeze. The claims catastrophe-response period runs from weeks 18 through 21; only low-risk migration work may occur then.\n- Available team: one program manager, two identity architects, four identity engineers, three endpoint engineers, two messaging engineers, three application engineers, two security analysts, one compliance lead, and part-time representatives from the service desk, HR, legal, and each business unit.\n- Approved tooling budget beyond existing licenses is $180,000. A specialist migration vendor may be engaged, but vendor selection and onboarding require four weeks.\n\nNon-negotiable controls\n1. No bulk migration wave may begin until identity correlation reaches at least 99.95% accuracy and every unresolved collision has a named owner.\n2. Every privileged identity must use phishing-resistant MFA before it enters the destination tenant.\n3. Application owners must sign off on authentication, authorization, provisioning, deprovisioning, logging, and rollback tests.\n4. Legal holds and retention labels must be validated before mailbox or SharePoint migration.\n5. The source tenant must remain recoverable and read-only for at least 90 days after the final user wave.\n6. Rollback must be defined in operational terms, including the latest safe decision point, responsible role, expected recovery time, and treatment of writes made after cutover.\n\nCreate a decision-ready execution plan in Markdown. It must be specific enough for teams to schedule and operate from, not a generic identity-migration overview. Include:\n\n1. Executive recommendation: the proposed migration pattern, why it is preferable to plausible alternatives, the five largest risks, and the decisions executives must make within the next two weeks.\n2. Current-state problem decomposition: identity authority, authentication, authorization, lifecycle management, applications, messaging and collaboration, endpoints, Azure resources, privileged access, compliance, and support readiness.\n3. Target-state architecture: authoritative sources, joiner/mover/leaver flow, account-correlation keys, domain and namespace handling, cross-tenant coexistence, Conditional Access, passwordless authentication, PIM, emergency access, guest identities, logging, and decommissioning boundaries. Clearly distinguish temporary coexistence controls from the permanent design.\n4. A 28-week phase plan that accounts for both change freezes. Show phases, entry and exit criteria, dependencies, accountable roles, staffing pressure points, and tangible deliverables. Include pilot cohorts and explain why they are representative.\n5. Migration-wave design: propose cohort sizes and sequencing for employees, contractors, privileged users, shared accounts, kiosks, and each major device class. Define go/no-go thresholds and a 24-hour, 72-hour, and 7-day stabilization process for every wave.\n6. Application disposition: classify the 73 federated applications, 18 SCIM integrations, 34 LDAP applications, and seven object-ID-dependent applications into migration patterns. Provide a practical strategy for the three applications that cannot bulk-remap users, including trade-offs and escalation triggers. Do not assume an undocumented vendor capability.\n7. Identity-correlation and collision-resolution procedure: matching hierarchy, confidence levels, treatment of dual identities, duplicate aliases, guests, contractors, shared mailboxes, and people with name changes. Include a control-total approach and a sample exception-record format.\n8. Security and compliance control matrix with control objective, implementation, evidence, owner, validation point, and failure response. Cover privileged access, phishing-resistant MFA, separation of duties, audit retention, legal holds, data residency, consent, dormant accounts, emergency access, and source-tenant access during the 90-day recovery period.\n9. Operational cutover runbook for one representative 500-user wave. Give a time-ordered schedule from T-7 days through T+7 days, with checkpoints, communications, monitoring, reconciliation, rollback decision points, and named roles rather than individual names.\n10. Test and acceptance plan covering positive, negative, performance, security, lifecycle, application, device, coexistence, disaster-recovery, and rollback testing. Include measurable acceptance thresholds.\n11. Support and communications plan for executives, managers, end users, application owners, the service desk, security operations, and regional privacy stakeholders. Include service-desk staffing assumptions, escalation paths, and message timing, but do not draft lengthy email copy.\n12. RAID register containing at least 15 substantive items, each with probability, impact, leading indicator, mitigation, contingency, owner, and decision deadline.\n13. Budget allocation for the $180,000 incremental budget. Present a base case and a contingency case, state assumptions, and flag where staffing capacity rather than tooling is the binding constraint.\n14. Program scorecard with no more than 12 metrics. Define formula, data source, reporting cadence, target, warning threshold, and executive interpretation for each metric.\n15. A final one-page steering-committee checklist of approvals and unresolved decisions.\n\nUse concise prose, tables where they improve usability, and explicit recommendations. Identify assumptions and unknowns rather than silently inventing facts. Resolve tensions among schedule, security, reversibility, staffing, and user disruption. If the 28-week deadline is infeasible under the stated controls, say so clearly and present the smallest defensible scope or resource adjustment that would make it achievable."
    }
  ]
}
Response
{
  "taskType": "textInference",
  "taskUUID": "adbd5955-f339-4262-b369-1abc2788a08e",
  "text": "# Entra ID Tenant Consolidation: Northstar Mutual ← Harbor Specialty\n**Program Execution Plan — Version 1.0 (Decision-Ready Draft)**\nPrepared by: Principal Identity Architect\nAudience: CIO, CISO, service owners, migration team, steering committee\n\n---\n\n## Feasibility Statement (Read First)\n\nThe 28-week deadline is **achievable only under a scoped-migration pattern with two conditions**:\n\n1. The **34 LDAP applications** are **not** re-platformed inside the program window. They remain on a coexistence Active Directory footprint (either a retained Harbor AD forest trusted from Northstar, or a new resource forest) and are re-platformed on a follow-on 6–9 month track.\n2. A **specialist migration vendor is engaged in Week 1–4** for tooling (cross-tenant sync accelerator, mailbox and OneDrive movement, correlation reporting). Without vendor tooling, in-house staffing is the binding constraint (see §13).\n\nIf executives insist on LDAP re-platforming inside 28 weeks, the defensible options are (a) extend to 40 weeks, (b) add three application engineers and one identity architect, or (c) descope 12 Azure subscription migrations to a follow-on wave. This document assumes the scoped pattern.\n\n---\n\n## 1. Executive Recommendation\n\n### Proposed pattern: **Staged coexistence-first consolidation into Tenant N**, using cross-tenant synchronization (CTS), progressive mail routing, and wave-based user cutover with source-tenant retained read-only for 90 days.\n\n| Element | Choice |\n|---|---|\n| Destination | Tenant N (cloud-native, PIM-mature, passwordless-forward) |\n| Coexistence engine | Entra Cross-Tenant Synchronization (B2B members) + mail coexistence via smart host |\n| Directory synchronization | Retain Harbor AD Connect during coexistence; cutover to Tenant N-anchored Connect for the two Harbor forests; groom out post-migration |\n| Application handling | Re-federate SAML/OIDC to Tenant N; SCIM re-point; LDAP defer; object-ID apps handled per §6 |\n| Data movement | Native cross-tenant mailbox migration + SPMT/Mover for SharePoint/OneDrive; Teams via cross-tenant Teams migration |\n| Devices | Intune tenant-to-tenant using Windows Autopilot re-enrollment for Windows; ADE re-enrollment for macOS; MAM re-enrollment for mobile |\n| Contractor devices | B2B guest with device compliance enforced via CA; no MDM enrollment |\n\n### Why this over plausible alternatives\n\n| Alternative | Why rejected |\n|---|---|\n| Big-bang cutover in one weekend | 4,600 employees + 900 contractors + 137 apps + 12 subscriptions exceeds recoverable blast radius; violates 15-minute SLA. |\n| Keep both tenants indefinitely with B2B only | Fails consolidation mandate, does not resolve privileged sprawl, sustains SMS-MFA debt, does not reduce license spend. |\n| Migrate to a **new** third tenant | Doubles work; abandons Tenant N's mature PIM, CA, and passwordless posture. |\n| Rebuild Harbor users in Tenant N from HR feed and abandon source | Loses mailbox history, breaks legal holds and retention labels, breaks object-ID dependent apps beyond repair. |\n\n### Five largest risks (full register in §12)\n1. **Object-ID-anchored apps without vendor remap** — three apps with no bulk remap; per-user re-linking at 5,500 users is infeasible without downtime concessions or bridging identity mapping.\n2. **Privileged-access exposure during coexistence** — 39 GAs, 11 standing, 6 shared emergency accounts in tenant H create a large attack surface until harmonized.\n3. **Correlation accuracy below 99.95%** — 1,200 guest collisions, 87 duplicate proxies, 46 dual identities; correlation gate directly controls schedule.\n4. **Mail routing outage during proxy-address hand-off** — customer-facing 15-minute SLA is unforgiving; misconfigured accepted domains or MX cutover causes hours of NDRs.\n5. **Change-freeze compression** — quarter-end (W9–11) and catastrophe response (W18–21) remove ~7 productive weeks; correlation, pilots, and final waves must fit around them.\n\n### Executive decisions required within 14 days\n\n| # | Decision | Owner | Trigger if not made |\n|---|---|---|---|\n| D1 | Approve scoped pattern (LDAP deferred) or fund extension | CIO | 28-week plan slips |\n| D2 | Approve vendor engagement and $180K envelope | CIO + CFO | Correlation and tooling delayed; slip ≥4 weeks |\n| D3 | Approve retirement of 6 Harbor shared emergency accounts and mandate PIM+FIDO2 for all admins | CISO | Non-negotiable Control 2 unmet |\n| D4 | Confirm Tenant N as authoritative and endorse `harborspecialty.com` addition as accepted domain | CIO + Legal | Namespace and email strategy blocked |\n| D5 | Legal/Privacy officers confirm cross-border processing model (US/UK/DE) and legal-hold custodianship transfer approach | Compliance Lead + Counsel | Control 4 unmet; mailbox migration blocked |\n| D6 | Approve contractor identity model (B2B members vs. guests) and device policy for 760 BYOD | CISO | Contractor cutover blocked |\n| D7 | Approve freeze exceptions policy and go/no-go authority (single named role) | CIO | Waves stall in freeze weeks |\n\n---\n\n## 2. Current-State Problem Decomposition\n\n| Domain | Tenant N | Tenant H | Gap / Risk |\n|---|---|---|---|\n| Identity authority | Cloud-first, HRIS-driven | Two on-prem AD forests → Entra Connect | Two forests must be reconciled; ownership of source-of-truth per attribute unclear |\n| Authentication | 68% passwordless (FIDO2/WHfB); CA baseline | Password + SMS MFA dominant | SMS MFA is not phishing-resistant; must uplift before privileged cutover |\n| Authorization | Group-based, PIM for privileged | Ad hoc; 39 GAs; 11 standing | Standing privilege violates least privilege |\n| Lifecycle (JML) | HR-driven, automated | Manual + scripted per forest | JML gaps produce dormant accounts; contractors unmanaged at scale |\n| Applications | Consolidated SSO catalog | 73 SAML/OIDC, 18 SCIM, 34 LDAP, 7 object-ID-dependent | LDAP apps block full decommission; object-ID apps threaten users |\n| Messaging & collab | Single EXO org | Separate EXO; overlapping proxies (87 duplicates) | Proxy collisions block mail routing |\n| Endpoints | Intune-managed, Autopilot | Intune, mixed enrollment maturity | Kiosks and BYOD contractors need bespoke approach |\n| Azure resources | Established landing zones | 12 subscriptions, unknown RBAC hygiene | Subscription transfer requires billing + RBAC redesign |\n| Privileged access | PIM, 2 monitored break-glass | 39 GA, 11 standing, 6 shared emergency | Immediate risk; cannot enter Tenant N as-is |\n| Compliance | Retention labels, legal holds mapped | Unknown label parity, hold inventory needed | Control 4 blocker |\n| Support readiness | Mature service desk | Separate ITSM, separate runbooks | Merged runbook + bilingual (EN/DE) support required |\n\n**Unknowns to close in Weeks 1–3:** exact legal-hold custodian list; Harbor SCIM app tokens and rotation authority; Harbor conditional access baseline; Azure subscription commercial ownership (EA vs. MCA); DE works-council notification status.\n\n---\n\n## 3. Target-State Architecture\n\n### Permanent design\n- **Authoritative sources:** HRIS (Workday) → Tenant N Entra ID for employees. Contractor management system → Tenant N with sponsor attribute mandatory. Application ownership records → ServiceNow CMDB.\n- **JML:** Workday event → Entra provisioning → downstream SCIM apps → Exchange, Teams, SharePoint, Intune. Leaver: 24-hour disable, 30-day soft delete, 7-year archive per policy (compliance to confirm).\n- **Correlation keys (in priority order):** `employeeId` (HRIS) → `mail` (primary SMTP normalized) → `onPremisesImmutableId` → `userPrincipalName` → `displayName + hireDate` (advisory only, never auto-merge). See §7.\n- **Namespaces:** `harborspecialty.com` and other Harbor vanity domains added to Tenant N as accepted domains and verified. UPN policy: Northstar keeps `@northstar.com`; migrated Harbor users retain existing UPN suffix. All Harbor SMTP addresses retained as proxy addresses.\n- **Conditional Access baseline:** device compliance + phishing-resistant MFA for all users; risk-based sign-in policy; location-aware controls for US/UK/DE; separate policy set for kiosks (shared device sign-in, short session, no persistent tokens); separate policy for contractors (MAM-WE, block non-compliant clients).\n- **Passwordless:** FIDO2 security keys for admins; Windows Hello for Business + Microsoft Authenticator (device-bound passkeys) for employees; extend to Harbor population as part of wave prerequisites.\n- **PIM:** All privileged roles time-bound, approval-required, MFA-on-activation. Zero standing GA except two monitored break-glass. Access reviews quarterly.\n- **Emergency access:** Two FIDO2-protected, monitored break-glass accounts, stored in split-knowledge safe. Alert on any sign-in.\n- **Guest identities:** B2B collaboration for external partners only; cross-tenant access settings restricted to trusted tenants; entitlement management for guest lifecycle.\n- **Logging:** Sign-in and audit logs to Log Analytics, 2-year hot, 7-year archive per compliance direction; Sentinel analytics for privileged actions and impossible travel.\n- **Decommissioning boundaries:** Tenant H set to read-only 30 days after final wave; retained 90 days (Control 5); formal tenant closure at Day 120; retention exports handed to compliance before closure.\n\n### Temporary coexistence controls (removed post-cutover)\n- Cross-tenant sync (Tenant H → Tenant N) creating B2B member objects.\n- Mail coexistence via shared namespace configuration and centralized mail transport (Tenant H accepted as internal relay for `harborspecialty.com` during transition).\n- Dual SSO for the 3 object-ID-dependent apps (see §6) with identity bridge.\n- Retained Harbor AD Connect until final Harbor user is cut over.\n- Cross-tenant access settings allowing seamless Teams meetings and file access.\n- Temporary CA policy allowing SMS MFA for non-privileged Harbor users during first 6 weeks after their cutover, with hard deadline for uplift.\n\n---\n\n## 4. 28-Week Phase Plan\n\nAssumes program start = Week 0. Freezes: W9–W11 (quarter-end), W18–W21 (catastrophe response).\n\n| Phase | Weeks | Entry criteria | Exit criteria | Key deliverables | Accountable | Pressure points |\n|---|---|---|---|---|---|---|\n| **P0 Mobilize** | 0–2 | Charter signed | RACI, vendor SOW draft, discovery kickoff | Program charter, RAID v1, comms plan | Program Manager | Vendor procurement lead time |\n| **P1 Discover & Baseline** | 1–5 | P0 exit | Full inventory: users, groups, apps, mailboxes, devices, subscriptions, legal holds; correlation dataset v1 | Discovery report, correlation model, hold inventory | Identity Architects | Legal-hold discovery latency |\n| **P2 Design & Vendor** | 3–6 | Discovery ≥80% complete | HLD/LLD signed; vendor onboarded; tooling deployed in lab | Target architecture, CA design, migration tooling PoC | Identity Architects + Vendor | Vendor onboarding 4-week runway |\n| **P3 Foundation Build** | 5–8 | Design signed | CTS operational; accepted domains verified; PIM extended; break-glass tested; correlation ≥99.5% | Foundation runbook, CA policies deployed in report-only | Identity Engineers | Domain verification, DNS ownership |\n| **⛔ Freeze 1** | 9–11 | — | No production identity changes | Regression tests, documentation, pilot prep | PM | Team morale; use for training |\n| **P4 Pilot Wave** | 12–14 | Correlation ≥99.95% (Control 1); privileged uplift complete (Control 2) | Pilot cohort migrated with zero Sev-1; app owner sign-offs (Control 3) | Pilot report, tuned runbook | Identity Architects | Cohort composition |\n| **P5 Privileged & Admin Migration** | 13–15 | Pilot pass; FIDO2 issuance complete for admins | All Harbor admins in Tenant N with PIM; standing GA = 0; shared emergency accounts retired | Admin migration report, PIM baseline | CISO delegate | Executive scheduling |\n| **P6 Wave 1 (Corporate Non-Customer-Facing)** | 15–17 | P5 exit; legal-hold validation for cohort (Control 4) | 800 users migrated; stabilization thresholds met | Wave 1 closure report | Migration Lead | Freeze 2 approaching |\n| **🟡 Restricted Window** | 18–21 | Catastrophe response | Only low-risk work: docs, contractor onboarding to B2B, app owner testing, non-production Azure subscription moves | Application readiness backlog burn-down | Migration Lead | Claims platform sacred; no customer-facing waves |\n| **P7 Wave 2 (Customer-Facing incl. Claims)** | 22–24 | P6 closure; claims platform coexistence test passed | 2,200 users; 15-minute SLA maintained; kiosk cohort migrated | Wave 2 report | Migration Lead + Business Unit reps | Telephony dependency |\n| **P8 Wave 3 (Remaining Employees + Contractors)** | 24–26 | P7 stabilization ≥72h clean | Remaining 1,600 employees + 900 contractors; BYOD onboarded | Wave 3 report | Migration Lead | Contractor identity provisioning |\n| **P9 Azure + Residual Apps** | 25–27 | Users migrated | 12 subscriptions transferred; residual app cleanup; source tenant read-only | Subscription transfer records | Cloud lead (identity architect) | Billing agreements |\n| **P10 Stabilize & Handover** | 27–28 | All waves complete | Ops runbook handover; scorecard closed; source tenant enters 90-day retention | Program closure report | PM | Documentation completeness |\n\n### Pilot cohort composition (155 users, Week 12)\nChosen to exercise every migration pattern in a controlled sample:\n- 25 IT staff (self-support, tolerant of issues)\n- 20 knowledge workers, US (baseline)\n- 20 knowledge workers, UK (regional CA)\n- 15 knowledge workers, DE (data-residency, works-council touch)\n- 15 claims adjusters, non-customer-facing shift (kiosk exposure)\n- 10 users of one object-ID-dependent app (Tier-A)\n- 10 users of an SCIM app + a SAML app\n- 10 users with dual identity in both tenants (§7)\n- 10 users with duplicate proxy addresses\n- 10 shared mailbox delegates\n- 5 executives (visibility)\n- 5 contractors (B2B pattern)\n\n**Representativeness rationale:** every collision class, every device class except macOS-heavy, every app class, both freeze-adjacent regions, and 3 language/regulatory contexts are exercised.\n\n### Staffing pressure points\n- **Weeks 5–8:** identity engineers over-allocated (foundation + design refinement). Vendor absorbs 40%.\n- **Weeks 12–17:** application engineers are the bottleneck (per-app cutover testing).\n- **Weeks 22–24:** service desk peak; contract 8 additional Tier-1 agents for 4 weeks.\n- **Weeks 25–27:** endpoint engineers on kiosk and macOS re-enrollment; consider deferring macOS beyond W28 if slipping.\n\n---\n\n## 5. Migration-Wave Design\n\n### Cohort sequencing\n\n| Wave | Weeks | Population | Size | Rationale |\n|---|---|---|---|---|\n| Pilot | 12–14 | Mixed representative | 155 | Prove runbook |\n| Admin | 13–15 | Privileged & shared accounts | ~90 | Control 2; must precede business waves |\n| W1 | 15–17 | Corporate back-office (finance, HR, legal, IT-remainder) | 800 | Low customer impact, high internal visibility |\n| Restricted | 18–21 | Contractors onboarded to B2B only; no cutover | 900 (staged) | Freeze compliance |\n| W2a | 22 | Claims adjusters non-kiosk | 900 | Business-critical, in-hours SLA |\n| W2b | 23 | Claims adjusters on kiosks | 480 kiosk sessions | Kiosk pattern; scheduled per shift |\n| W2c | 24 | Customer service + telephony users | 820 | Telephony coordination |\n| W3a | 25 | Remaining employees | 1,000 | Long tail |\n| W3b | 26 | Remaining employees + contractors | 600 + 900 | Cleanup |\n\n### Device class treatment\n- **Windows laptops (7,900 total; ~3,100 Harbor):** Autopilot re-enrollment using ESP; user data preserved via OneDrive Known Folder Move confirmed before cutover.\n- **macOS (2,300 total; ~900 Harbor):** ADE re-enrollment; Platform SSO with new tenant; scheduled during user off-hours.\n- **Managed mobile (3,100):** Company Portal re-enroll; MAM policy replacement.\n- **Kiosks (480):** Shared-device mode; wave in one facility per night; local support present.\n- **Contractor BYOD (760):** No MDM; B2B guest or member with MAM-WE and CA compliance grant.\n\n### Go/no-go thresholds (each wave)\n\n| Threshold | Value | Action if breached |\n|---|---|---|\n| Correlation accuracy for cohort | ≥99.95% | Delay wave |\n| Legal-hold validation completeness | 100% | Delay wave |\n| App owner sign-off for cohort's apps | 100% of Tier-A, ≥95% Tier-B | Delay wave |\n| Privileged users on FIDO2 | 100% of cohort admins | Delay wave |\n| Pilot Sev-1 defects open | 0 | Delay wave |\n| Service desk staffing | ≥120% of forecast | Delay wave |\n| Freeze exception approved (if applicable) | Written approval | Cancel wave |\n\n### Stabilization protocol per wave\n\n| Interval | Actions | Owner | Exit |\n|---|---|---|---|\n| T+0 to T+24h | Hyper-care bridge open; sign-in success rate monitored per 15 min; mail flow test batches every hour; top-10 error triage | Migration Lead + Service Desk | <2% ticket-per-user rate; no Sev-1 |\n| T+24h to T+72h | Daily reconciliation report (accounts, licenses, group memberships, mailbox status); app usage anomaly check | Identity Engineer on-call | All reconciliation deltas explained |\n| T+72h to T+7d | Delegated permissions verification; retention label validation; guest access smoke test; wave retrospective | Migration Lead | Retrospective actions logged; wave closure signed by CISO delegate |\n\nRollback decision points defined per wave in §9.\n\n---\n\n## 6. Application Disposition\n\n### Classification\n\n| Class | Count | Pattern | Effort |\n|---|---|---|---|\n| SAML/OIDC re-federation, low complexity | ~50 of 73 | Add Tenant N as second IdP → cutover metadata → decommission Tenant H app | S |\n| SAML/OIDC with claim/attribute customization | ~20 of 73 | As above + claims re-mapping; regression test | M |\n| SAML/OIDC with hard IdP entity ID lock | ~3 of 73 | Vendor ticket to change entity ID; may need re-registration; schedule per-app cutover | L |\n| SCIM provisioning | 18 | Re-point SCIM endpoint to Tenant N; rotate tokens; scope-test; run parallel provisioning for 48h; then disable Tenant H | M |\n| LDAP via on-prem AD | 34 | **Defer.** Retain Harbor AD forest (or replace with resource forest) trusted from Northstar. Re-platform on 6–9 month follow-on. | Out of scope |\n| Object-ID dependent, vendor-supported remap | 4 of 7 | Export mapping (old objectId → new objectId), run vendor API remap during cutover window, verify | M |\n| Object-ID dependent, **no vendor remap** | 3 of 7 | See below | H |\n| Azure subscriptions | 12 | Transfer to Tenant N EA/MCA; re-do RBAC from group entitlements (do not migrate identities as-is) | M |\n\n### Three apps without bulk remap — practical strategy\n\nAssumption: no undocumented vendor capability exists. Options in decreasing order of preference:\n\n**Option A — Identity Bridge (recommended for Tier-A):**\n- Keep app federated to Tenant H IdP during coexistence; add Tenant N users as B2B guests to Tenant H for authentication only.\n- User experience: single sign-on via Tenant N credentials proxied through B2B; app sees original Tenant H `objectId`.\n- Trade-off: preserves object-ID; delays true consolidation of that app to a follow-on project; requires Tenant H to remain live for authentication (adjust Control 5 to authentication-only, not full read-only).\n- Escalation trigger: if Tenant H cannot remain live for authentication (e.g., licensing cost, security posture), move to Option B.\n\n**Option B — Manual per-user re-linking with scheduled downtime:**\n- Coordinate app owner + vendor for maintenance window; export user list; on cutover, admin creates a new record in the app for each Tenant N identity; historical data re-associated via a mapping table maintained in the app database (requires vendor read/write DB agreement).\n- Trade-off: downtime per user or per batch; risk of lost historical linkage; vendor may not warrant DB manipulation.\n- Escalation trigger: any vendor refusal → Option C.\n\n**Option C — Replace the application:**\n- Time-boxed evaluation (30 days) of alternatives; if replacement identified, put app in the 6–9 month follow-on backlog and continue coexistence via Option A.\n\n**Recommendation:** Apply Option A for all three during the 28-week window; commit to a Q+2 project to eliminate the bridge. This preserves the schedule and Control 1 without pretending vendor capability exists.\n\n---\n\n## 7. Identity Correlation & Collision Resolution\n\n### Matching hierarchy\n\n| Rank | Key | Confidence | Action |\n|---|---|---|---|\n| 1 | `employeeId` present in both tenants and equal | High | Auto-correlate |\n| 2 | Primary SMTP normalized match AND HR record match | High | Auto-correlate |\n| 3 | `onPremisesImmutableId` match | High | Auto-correlate |\n| 4 | UPN local part + hire date + manager match | Medium | Auto-correlate with sampling audit |\n| 5 | `displayName + department + country` match | Low | Route to exception queue; human decision |\n| 6 | Fuzzy match (Jaro-Winkler ≥0.92 on name + shared attributes) | Advisory only | Never auto-merge; suggest to reviewer |\n\n### Category treatment\n\n| Category | Volume | Treatment |\n|---|---|---|\n| Dual identity (person in both) | 46 | Person-by-person interview via HR partner; retain Tenant N identity as primary, migrate Tenant H mailbox contents, add Harbor SMTP as alias, transfer group memberships explicitly |\n| Duplicate proxy addresses | 87 | Compare owners; if same person → merge; if different → keep Harbor SMTP on Harbor account only, rename Tenant N proxy per naming standard, notify affected users 14 days before cutover |\n| Duplicate display names | 640 | Automatic disambiguation using department suffix; no name change required; documented in comms |\n| Guest collisions | 1,200 | Delete stale guests (>90 days inactive) after owner sign-off; migrate active guests via CTS; re-invite where guest was a Harbor employee |\n| Contractors | 900 | Fresh B2B member creation in Tenant N with sponsor attribute; no correlation to Tenant N employees |\n| Shared mailboxes | Discovery pending | Recreate in Tenant N; migrate content; reapply delegate permissions from export; do not convert to user mailboxes |\n| Recent name changes | Discovery pending | Match on `employeeId` + previous UPN captured from Entra audit logs |\n\n### Control totals\nBefore every wave, publish reconciliation showing:\n`in-scope users` = `auto-correlated` + `manually resolved` + `explicit exceptions with owner`\nAny residual = wave blocked.\n\n### Exception record format\n\n```\nExceptionID: HAR-EX-0007\nType: DualIdentity | DuplicateProxy | GuestCollision | Other\nSourceObjectId (Tenant H): <guid>\nTargetObjectId (Tenant N): <guid or NEW>\nMatchConfidence: High | Medium | Low | None\nProposedResolution: Merge | KeepBoth | RenameProxy | ManualCreate\nOwner: <role, e.g., HR Business Partner - EMEA>\nDecisionDeadline: <YYYY-MM-DD>\nApprovedBy: <role>\nEvidenceLinks: [ticket, HR record, email]\nNotes:\n```\n\n---\n\n## 8. Security & Compliance Control Matrix\n\n| Objective | Implementation | Evidence | Owner | Validation point | Failure response |\n|---|---|---|---|---|---|\n| Least-privileged admin | PIM eligible-only; 0 standing GA except 2 break-glass | PIM export weekly | CISO delegate | Weekly and pre-wave | Suspend privileged access; incident |\n| Phishing-resistant MFA for admins | FIDO2 required for privileged; CA policy | CA policy export; sign-in logs | Identity Architect | Pre-P5 gate | Block admin migration |\n| Separation of duties | Migration engineers cannot approve their own change; PIM approver ≠ requester; wave sign-off by CISO delegate | Change tickets; PIM approvals | PM | Every change | Change reversed |\n| Audit log retention | Log Analytics 2y hot, 7y archive | Retention config screenshot | Security Analyst | Weekly | Escalate to CISO |\n| Legal holds validated pre-migration | Compliance queries mailbox and site holds; sign-off per user cohort | Hold inventory report + sign-off | Compliance Lead + Counsel | Pre-wave gate (Control 4) | Wave halted |\n| Data residency | Multi-geo EXO and SharePoint honored; user residency attribute drives placement | Multi-geo config; sample validation | Compliance Lead | Pre-cutover per region | User held back |\n| Consent (DE works council, contractor privacy) | Privacy officer approval on data flow; DPIA update | DPIA doc; approval email | Compliance Lead + Privacy Officer | Pre-P4 | Region scope delayed |\n| Dormant accounts | Disable accounts inactive >90d before migration | Inactivity report | Identity Engineer | Pre-wave | Excluded from wave |\n| Emergency access | 2 break-glass FIDO2, split-knowledge, alert on sign-in | Vault log; alert config | CISO delegate | Monthly test | Rotate immediately |\n| Source-tenant 90-day period | Tenant H set read-only; only 2 named admins retain access via PIM; all sign-ins alerted | Tenant H CA config; sign-in logs | Identity Architect | Continuous | Freeze access; investigate |\n| Change control | All migration changes via CAB-lite with pre-approved templates | CAB records | PM | Weekly | Rollback |\n| SMS MFA sunset | Temporary 6-week window post-cutover per user; forced enrollment to Authenticator/FIDO2 | Auth method report | Identity Engineer | 6 weeks post-wave | Account MFA reset |\n\n**Matters requiring counsel or privacy-officer approval (not decided here):**\n- Custodianship of legal holds when mailboxes move between tenants.\n- Cross-border processing of German employee data during coexistence.\n- Works-council consultation requirements in Germany.\n- Contractor personal-device data-processing scope for B2B members vs. guests.\n- Retention obligations for the Tenant H tenant itself post-decommission.\n- Records-retention treatment of correlation exception decisions.\n\n---\n\n## 9. Cutover Runbook — Representative 500-User Wave\n\nTimes are in the wave's primary business time zone. Roles are named by function, not person.\n\n### T-7 days\n- Comms: Manager-to-user reminder issued (Service Desk Comms Lead).\n- Correlation report re-run for cohort; must be 100% resolved (Identity Engineer).\n- Legal-hold validation confirmed for cohort mailboxes (Compliance Lead).\n- App owner sign-offs collected for all Tier-A and Tier-B apps in cohort scope (App Engineering Lead).\n- Device readiness report: OneDrive KFM confirmed for each Windows user (Endpoint Lead).\n\n### T-3 days\n- Freeze secondary changes (Change Manager).\n- Hypercare bridge scheduled; on-call rota confirmed (Migration Lead).\n- Reversal test: verify last known-good backup of source objects (Identity Engineer).\n- Executive brief sent (PM).\n\n### T-1 day\n- Pre-flight checklist reviewed live in go/no-go call at 14:00 local (Migration Lead chairs).\n- **Latest safe decision point: T-1 day 18:00.** If aborted after this, communications sent within 1 hour.\n- Break-glass account tested (CISO delegate).\n- Service desk pre-brief and macro loading (Service Desk Manager).\n\n### T-0 (cutover Friday 20:00 → Saturday 06:00)\n- 20:00 Freeze source-tenant provisioning (Identity Engineer).\n- 20:15 Disable password writeback on Harbor Connect (Identity Engineer).\n- 20:30 Trigger CTS delta run; verify object counts (Identity Engineer).\n- 21:00 Mailbox move requests submitted in batches of 50 (Messaging Engineer).\n- 22:00 SCIM re-point for cohort apps (App Engineer).\n- 23:00 SAML metadata cutover for cohort apps (App Engineer).\n- 00:00 Intune re-enroll batch commences (Endpoint Engineer).\n- 03:00 Smoke test suite executes (Test Lead).\n- 04:00 Mail routing update for cohort's proxy addresses (Messaging Engineer).\n- 05:00 Reconciliation control totals published (Identity Engineer).\n- 05:30 Go/no-go for opening business (Migration Lead + CISO delegate + Business Unit rep).\n- 06:00 Hypercare bridge remains open; business go-live.\n\n**Rollback decision points (T-0):**\n| Time | Trigger | Decision role | Recovery time | Post-cutover writes handling |\n|---|---|---|---|---|\n| 22:30 | Mailbox move failure rate >5% | Migration Lead | 2h to source state | None yet; no data loss |\n| 00:30 | SSO failure for >2 apps | Migration Lead | 3h to source | None yet |\n| 03:30 | Smoke test critical fail | Migration Lead + CISO delegate | 4h; users remain in Tenant H Monday | None yet |\n| 05:30 | Reconciliation deltas unexplained >0.05% | Migration Lead + CISO delegate | 4h | None yet |\n| **Post 06:00** | Sev-1 emerges after go-live | CIO delegate | Rollback becomes forward-fix; roll back only individual users; source-tenant writes captured via delta report and replayed manually | Written to Tenant N; source is read-only; per-user rollback recreates account and replays via export |\n\n### T+24h\n- Reconciliation report v2 (Identity Engineer).\n- Ticket triage summary (Service Desk Manager).\n- Sev-1/Sev-2 review at 10:00 (Migration Lead).\n\n### T+72h\n- Delegated permissions verification (Messaging Engineer).\n- Group membership audit (Identity Engineer).\n- App owner post-cutover confirmation (App Engineer).\n\n### T+7d\n- Wave retrospective (PM).\n- Formal closure sign-off (CISO delegate).\n- Lessons applied to next wave runbook (Migration Lead).\n\n---\n\n## 10. Test & Acceptance Plan\n\n| Test class | Scope | Acceptance threshold |\n|---|---|---|\n| Positive functional | Sign-in, mail send/receive, Teams chat/meet/call, SPO access, Intune app deploy | 100% pass on smoke suite |\n| Negative | Blocked-country sign-in, non-compliant device, expired MFA method | 100% blocked as designed |\n| Performance | Sign-in latency, mailbox move throughput, provisioning latency | Sign-in p95 <3s; provisioning <15 min end-to-end |\n| Security | Privileged access review, CA bypass attempt, token replay, break-glass alert | 0 successful bypass; 100% alert fire |\n| Lifecycle (JML) | New hire, transfer, termination in migrated population | Provisioning ≤4h; termination ≤1h |\n| Application | Per-app: authN, authZ, provisioning, deprovisioning, logging, rollback (Control 3) | 100% owner sign-off before wave |\n| Device | Autopilot, ADE, mobile enroll, kiosk shared-mode | ≥98% success on first attempt |\n| Coexistence | Cross-tenant Teams, calendar free/busy, mail flow both directions | 100% pass on 25-scenario matrix |\n| DR | Simulated tenant-level auth outage; break-glass exercised | Access restored <30 min |\n| Rollback | Per-wave dry run in lab; per-user rollback in production | Dry run: full source restore <4h; per-user: <1h |\n\nTest evidence is filed against wave record; missing evidence = no wave sign-off.\n\n---\n\n## 11. Support & Communications Plan\n\n### Audience matrix\n\n| Audience | Message | Cadence | Owner |\n|---|---|---|---|\n| Executives | Scorecard, risk deltas, decisions needed | Weekly (60-min) | PM |\n| Managers | Wave schedule for their teams, actions required | 4, 2, 1 weeks pre-wave | Business Unit reps |\n| End users | What is changing, when, what to do | 14, 7, 1 day pre-wave; day-of; +2 days | Service Desk Comms Lead |\n| Application owners | Test schedule, sign-off requests, cutover role | Continuous per app | App Engineering Lead |\n| Service desk | Runbooks, macros, escalation | Weekly plus pre-wave brief | Service Desk Manager |\n| SecOps | Detection tuning, wave anomalies expected | Pre-wave and daily during hypercare | Security Analyst |\n| Regional privacy stakeholders (US/UK/DE) | Data-flow diagrams, DPIA updates, works-council notes | Milestone-based | Compliance Lead |\n| Legal | Hold custodianship, tenant retention, contract items | Milestone-based | Compliance Lead |\n\n### Service desk staffing assumption\n- Baseline capacity handles 2,000 tickets/week.\n- Wave surge assumption: 12% of migrated users create a ticket; 3% Sev-2+; peak day at T+1 business day.\n- For W2 (2,200 users): +8 Tier-1 agents for 4 weeks, +2 Tier-2 for 6 weeks. German-language coverage required during DE cohort waves.\n\n### Escalation path\nSev-3 → Tier-1 → Tier-2.\nSev-2 → Migration Lead within 30 min.\nSev-1 → Migration Lead + CISO delegate + PM within 15 min; CIO in loop within 30 min; hypercare bridge activated.\n\n### Message timing (subject-line themes; drafts owned by Comms)\n- T-14d: \"What is changing and when\"\n- T-7d: \"Your action items before Friday\"\n- T-1d: \"Tonight's timeline and support numbers\"\n- T+0: \"You are now on the new platform — first steps\"\n- T+2d: \"Common questions and fixes\"\n\n---\n\n## 12. RAID Register\n\nR = Risk, A = Assumption, I = Issue, D = Dependency. Probability (P) and Impact (I) on 1–5.\n\n| # | Type | Item | P | I | Leading indicator | Mitigation | Contingency | Owner | Decision deadline |\n|---|---|---|---|---|---|---|---|---|---|\n| 1 | R | 3 object-ID apps with no bulk remap | 5 | 5 | Vendor confirms no API | Option A identity bridge | Option B downtime; Option C replace | App Eng Lead | W3 |\n| 2 | R | Privileged sprawl in Tenant H (39 GA) | 4 | 5 | GA audit >2 | Enforce PIM, retire shared accounts | Emergency access revocation | CISO delegate | W2 |\n| 3 | R | Correlation misses 99.95% | 4 | 5 | Weekly correlation trend | HR partner desk; exception SLAs | Slip wave 2 weeks | Identity Architect | Rolling |\n| 4 | R | Mail-flow outage at cutover | 3 | 5 | Failed dry runs | Rehearsals; smart-host; fallback MX | Revert mail routing (<30 min) | Messaging Eng | Pre-wave |\n| 5 | R | Freeze exception denied for critical fix | 3 | 4 | CAB backlog | Pre-approved emergency template | Halt wave | PM | W8 |\n| 6 | R | Vendor onboarding slippage | 3 | 4 | Contract redlines W3 | Parallel two-vendor evaluation | In-house scripts for CTS + reporting | PM | W2 |\n| 7 | R | Kiosk shared-mode conflict with claims platform | 3 | 4 | Pilot kiosk test fail | Isolated pilot facility | Delay kiosk cohort to post-W28 | Endpoint Lead | W12 |\n| 8 | R | DE works-council delay | 3 | 4 | Meeting minutes | Early engagement W2 | Descope DE from W2 wave | Compliance Lead | W6 |\n| 9 | R | 15-min SLA breach during W2 | 2 | 5 | Pilot recovery time >15 min | Pre-staged rollback; per-user cutover for claims | Individual rollback | Migration Lead | W20 |\n| 10 | R | Contractor B2B lifecycle unmanaged | 3 | 3 | No sponsor assigned in HR feed | Sponsor mandatory attribute; entitlement mgmt | Bulk disable at 30 days no sponsor | Identity Eng | W10 |\n| 11 | R | Legal-hold gap discovered mid-wave | 2 | 5 | Sample audit fails | Pre-wave 100% inventory (Control 4) | Halt affected mailbox moves | Compliance Lead | Pre-wave |\n| 12 | R | SharePoint permission drift post-migrate | 3 | 3 | Reconciliation deltas | Export/re-apply script; owner audit | Owner remediation SLA 5 days | App Eng Lead | Per-wave |\n| 13 | R | Azure subscription RBAC re-mapped incorrectly | 3 | 4 | Missing group entitlements | Pre-build target group model; parallel run | Rollback subscription move (billing action) | Cloud Lead | Pre-P9 |\n| 14 | R | Passwordless enrollment lag in Harbor | 4 | 3 | Enrollment <50% at W10 | Pop-up clinics; kit distribution | Extend temporary SMS window per user (not admin) | Identity Eng | W10 |\n| 15 | R | SecOps drowned in benign anomalies during waves | 3 | 3 | Alert volume >2× baseline | Pre-wave detection tuning; expected-anomaly runbook | Suppress specific rules with sunset date | Security Analyst | Pre-wave |\n| 16 | A | Harbor forests can remain during coexistence | — | — | AD forest health report | Confirm W1 | Retire on follow-on track | Identity Architect | W3 |\n| 17 | D | Vendor SOW signed by W4 | — | — | Contract status | Legal fast-track | In-house fallback plan | PM | W3 |\n| 18 | I | Six shared emergency accounts in Tenant H (Control 2) | 5 | 5 | Existing | Retire before W5; create individual PIM-eligible admins | Rotate credentials immediately if retention required | CISO delegate | W3 |\n\n---\n\n## 13. Budget Allocation\n\nIncremental envelope: **$180,000**.\n\n### Base case\n| Line | Cost | Rationale |\n|---|---|---|\n| Specialist migration vendor (tooling + advisory) | $95,000 | CTS accelerator, mailbox/SPO movement, correlation reporting, 12 weeks |\n| FIDO2 security keys (admins + high-risk roles, ~600 keys @ $55) | $33,000 | Control 2; buffer for loss |\n| Kiosk shared-mode licensing gap / device certs | $8,000 | 480 kiosks |\n| Correlation & reconciliation SaaS (identity governance add-on trial + entitlement mgmt tuning) | $12,000 | Exception queue tooling |\n| Third-party mailbox coexistence add-on (fallback licensing) | $10,000 | Only if native path insufficient in pilot |\n| Contingency reserve | $22,000 | 12% |\n| **Total** | **$180,000** | |\n\n### Contingency case (if vendor slips or 3 object-ID apps require Option B)\n- Move $22K reserve + $10K coexistence line into: additional application-engineering contractor hours ($20K), extended service-desk augmentation ($12K).\n- Trade-off: no third-party mail fallback; must rely on native path performing.\n\n### Assumptions\n- Existing Microsoft licenses cover Entra ID P2, Intune, Defender, Sentinel.\n- German-language service desk capacity is procurable within SLA.\n- Vendor pricing is within typical SOW ranges; a competitive selection is possible in 3 weeks.\n\n### Where staffing is the binding constraint (not money)\n- **Application engineering (3 FTE)** during W12–W17: even with unlimited budget, ramp-up time makes contract hires marginal. Recommend cross-training two identity engineers on SAML/SCIM basics in P0.\n- **Endpoint engineering (3 FTE)** during W22–W26: kiosk + macOS + Windows re-enrollment concurrent; consider deferring macOS cohort into follow-on.\n- **Compliance lead (1 FTE)** across all pre-wave gates: single point of failure for legal-hold sign-off. Nominate a backup with formal delegation from Counsel.\n\n---\n\n## 14. Program Scorecard\n\n| # | Metric | Formula | Source | Cadence | Target | Warning | Executive interpretation |\n|---|---|---|---|---|---|---|---|\n| 1 | Correlation accuracy | Auto+resolved / total in scope | Correlation tool | Weekly | ≥99.95% | <99.90% | Wave gate |\n| 2 | Privileged accounts on FIDO2 | FIDO2-enrolled admins / total admins | Entra auth methods | Weekly | 100% | <100% | Control 2 breach |\n| 3 | Standing GA count | Direct query | Entra PIM | Weekly | 2 (break-glass) | ≥3 | Immediate review |\n| 4 | Wave completion vs plan | Users migrated / plan | Migration tool | Weekly | 100% cumulative | −5% | Schedule risk |\n| 5 | Sev-1 incidents per wave | Count | ITSM | Per wave | 0 | ≥1 | Rollback consideration |\n| 6 | Business-hour outage minutes (customer-facing) | Cumulative minutes | ITSM | Per wave | 0 | ≥15 | SLA breach |\n| 7 | Ticket rate per 100 migrated users | Tickets / users × 100 | ITSM | Daily during hypercare | <12 | >18 | Support capacity |\n| 8 | Legal-hold validation coverage | Validated / migrated cohort | Compliance tool | Per wave | 100% | <100% | Wave halted |\n| 9 | App owner sign-off coverage | Signed / cohort apps | Program tracker | Per wave | 100% Tier-A; ≥95% Tier-B | Below | Wave halted |\n| 10 | Passwordless adoption (migrated Harbor) | Passwordless users / migrated | Auth methods | Weekly | ≥70% by W+6 post-wave | <50% | Adoption push |\n| 11 | Source-tenant sign-ins during 90d retention | Sign-in count from named admins only | Audit logs | Weekly | Only 2 named admins | Any other | Investigate |\n| 12 | Budget consumed vs plan | Spent / planned | Finance | Bi-weekly | ≤105% | >110% | Rebaseline |\n\n---\n\n## 15. Steering Committee One-Page Checklist\n\n**Approvals required at next steering meeting**\n- [ ] D1 — Endorse scoped migration pattern (LDAP deferred to follow-on)\n- [ ] D2 — Approve vendor selection and $180K budget envelope\n- [ ] D3 — Mandate retirement of Harbor shared emergency accounts and FIDO2 for admins by W5\n- [ ] D4 — Confirm Tenant N as destination; approve Harbor domain addition\n- [ ] D5 — Counsel and Privacy Officer confirmations for cross-border data flow and legal-hold custodianship\n- [ ] D6 — Approve contractor identity model and BYOD policy\n- [ ] D7 — Name single freeze-exception approver and go/no-go authority\n- [ ] Approve wave schedule respecting quarter-end (W9–W11) and catastrophe (W18–W21) freezes\n- [ ] Approve identity bridge (Option A) for three non-remappable applications with Q+2 exit plan\n- [ ] Approve 90-day Tenant H retention posture with 2 named admins\n\n**Unresolved decisions to schedule**\n- Follow-on program budget and timeline for 34 LDAP apps\n- Follow-on decision to eliminate identity bridge on 3 object-ID apps\n- Final DE works-council engagement scope and timing\n- macOS deferral trigger criteria if W25–W26 capacity is exceeded\n- Retention treatment of correlation exception records (records-retention policy interpretation → Counsel)\n- Post-program identity governance operating model (steady-state RACI)\n\n**Key dates to note**\n- W2: D1–D7 decisions due\n- W4: Vendor onboarding complete\n- W8: Foundation exit gate (pre-freeze)\n- W12: Correlation gate (Control 1) and pilot launch\n- W15: Privileged migration complete (Control 2)\n- W22–W24: Customer-facing wave (15-minute SLA in force)\n- W27: Source tenant enters 90-day retention\n- W+90 days after final wave: Tenant H closure decision\n\n---\n\n*End of plan. Prepared for decision, not for record. Update after each steering meeting; re-baseline scorecard at end of each phase.*",
  "cost": 0.438005,
  "finishReason": "stop"
}

reference-to-text

Textile Conservation Intake Report$0.078~48s
# Preliminary Condition Intake Report

## 1. Object Identification
**Object name:** Beaded frame purse (evening bag / reticule with kiss-lock closure).
**Summary:** A small rounded beaded bag with a metal frame and ball-clasp closure, featuring an all-over geometric beaded pattern in cream, brown, and dark red/purple tones, exhibiting multiple areas of loss, abrasion, and structural instability.

## 2. Visual Description
- **Form:** Rounded, roughly semi-circular to shield-shaped body suspended from a rigid metal top frame. The frame terminates in a central twin-ball kiss-lock clasp.
- **Frame:** Metal (tentatively a white metal or silvered alloy based on tone), with visible decorative stamped or chased ornament along the upper band and hinged pivot points at each end.
- **Beadwork:** Densely worked seed beads covering the entire visible face. The pattern comprises concentric lozenge/diamond motifs arranged in a lattice, alternating cream-white centers framed by golden-brown/amber and dark red-purple bead courses. A scalloped/undulating band of dark beads outlines each motif.
- **Additional components:** A circular tan-colored applied element (tentatively leather or hide) is present in the lower-right quadrant, from which a short frayed fibrous tuft/tassel remnant extends. Loose fibrous strands are also visible along the lower edge.

## 3. Condition Observations

**Structural fabric / ground textile**
- Exposed tan ground material visible at the upper-right shoulder of the bag where beadwork is lost, indicating localized separation of beads from substrate.
- Additional smaller exposed area at the upper-left curve of the frame juncture.
- Loose fibers protruding along the lower edge and lower-right, suggesting weakened seam threads or a degraded tassel/fringe attachment.

**Beadwork**
- Bead loss concentrated at the upper-right shoulder (largest loss area), with smaller losses at the upper-left near the frame.
- Scattered thread slackening tentatively indicated by slight irregularity in bead alignment along the lower-left curve.
- Overall bead surface appears intact in color and finish across the central field; no obvious large-scale iridescence loss discernible from the image.

**Metal frame and clasp**
- Frame surface shows uneven tonality and possible tarnish or patina, most visible along the upper edge and around the clasp.
- A darker discoloration/possible residue is visible on the inner face of the frame just left of the clasp (tentative: tarnish, adhesive residue, or accreted soil).
- Kiss-lock ball clasp appears present and aligned; functional status cannot be confirmed from the photograph.
- Hinge points at both frame termini are visible; integrity not verifiable from image alone.

**Surface soiling / staining**
- General overall dulling consistent with age-related soiling.
- The tan circular applied element in the lower-right shows tonal variation that may indicate staining, wear, or original patina (tentative).

**Deformation**
- Slight asymmetry between the left and right lower curves; whether this reflects original construction or deformation of the frame/textile is not determinable from a single view.
- No pronounced creasing or collapse evident in this view.

**Previous repairs**
- No unambiguous evidence of prior repair visible; the tan circular applied patch in the lower-right may represent either an original decorative element (e.g., tassel anchor) or a later intervention — this is **tentative** and requires in-person assessment.

## 4. Condition Priority Table

| # | Observation | Location | Priority | Justification |
|---|---|---|---|---|
| 1 | Bead loss with exposed substrate | Upper-right shoulder | **High** | Active loss zone; adjacent beads at risk of cascading detachment. |
| 2 | Secondary bead loss | Upper-left near frame | Moderate | Smaller area but in a stress zone adjacent to the rigid frame. |
| 3 | Loose/frayed fibers and tuft remnant | Lower-right and lower edge | Moderate | Indicates degraded thread or fringe attachment; vulnerable to snagging. |
| 4 | Possible thread slackening in beadwork | Lower-left curve | Moderate | Precursor to further bead loss if handled. |
| 5 | Metal frame tarnish / discoloration | Upper frame and area beside clasp | Low | Cosmetic; stable if kept in controlled RH. |
| 6 | Uncertain status of clasp/hinge function | Frame termini and central clasp | Moderate | Operability unknown; forcing risks damage. |
| 7 | Overall surface soiling | Entire face | Low | Non-active; addressed only after conservator assessment. |
| 8 | Possible minor asymmetry/deformation | Lower body outline | Low | Not confirmed; monitor. |

## 5. Handling, Packing, and Storage Recommendations
- Handle with clean nitrile gloves; support the object fully from beneath on a padded board — never lift by the frame alone, as the beaded body's weight may stress weakened attachment threads.
- Do **not** open or operate the kiss-lock clasp until examined by a conservator.
- Store flat, face-up, on an inert padded support (e.g., washed cotton stockinette over polyester batting on acid-free board), with a shaped ethafoam or rolled tissue support beneath the body to maintain its three-dimensional profile without pressure on the beadwork.
- Cover loosely with acid-free tissue; avoid direct contact between tissue and loose bead areas — use a raised tissue tent if needed.
- House in an acid-free box slightly larger than the object; label the exterior to indicate fragile beadwork and loose components.
- Maintain stable environment: ~18–20 °C, 45–55% RH, low light (<50 lux when on view); isolate from sulfur-containing materials to limit further metal tarnishing.
- Collect and retain any detached beads or fibers in a labeled archival vial or polyethylene bag stored with the object.

## 6. Recommended Examination and Documentation Steps

**Concluded from the photograph:**
- Presence and approximate location of bead loss zones.
- General pattern, palette, and construction type (beaded body on textile ground, metal kiss-lock frame).
- Presence of loose fibers and an applied circular element in the lower-right.
- Evidence of surface soiling and possible frame tarnish.

**Requires in-person inspection:**
- Verification of bead type, ground textile fiber, thread type, and lining materials.
- Frame metal identification (visual, and where appropriate, non-destructive analytical testing).
- Assessment of clasp and hinge mechanical function.
- Interior examination for lining condition, maker's marks, labels, or inscriptions.
- Reverse-side condition (not visible in this image).
- Confirmation of whether the tan circular element is original or a repair.
- Dimensions, weight, and detailed photographic documentation (raking light, macro of loss areas, UV examination for adhesives/repairs).
- Stitch/thread analysis under magnification to assess extent of thread degradation.

## 7. Collection-Management Summary
This beaded frame purse presents in fair but fragile condition. The beaded lattice pattern and metal kiss-lock frame remain visually coherent, but localized bead loss at the upper-right shoulder, secondary loss near the upper-left frame, and loose fibers along the lower edge indicate active vulnerability. Priority actions are stabilized housing, restricted handling, and referral to a textile conservator for full examination, substrate and thread assessment, and treatment planning. Provenance, date, materials, and maker are undetermined and require further research.

## 8. Alt Text
A small rounded beaded evening bag with a metal kiss-lock frame, photographed flat against a grey background beside a scale and color reference card. The bag is decorated with a geometric lattice of cream, amber-brown, and dark red-purple seed beads, and shows areas of bead loss at the upper right and loose fibers along the lower edge.
Finish reason: stop
Try in Playground
import { createClient } from '@runware/sdk'

const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()

const [result] = await client.run({
  model: 'anthropic:claude@opus-4.7',
  settings: {
    thinkingLevel: 'adaptive'
  },
  inputs: {
    images: [
      'https://assets.runware.ai/assets/inputs/019c3efd-2553-405e-ba95-0ac635f2f2bc.jpg'
    ]
  },
  messages: [
    {
      role: 'user',
      content: 'Inspect the attached conservation photograph and draft a preliminary museum condition intake report for the evening bag. Ground every statement in visible evidence and do not invent provenance, dimensions, maker, date, materials, construction methods, or damage that cannot be supported by the image.\n\nStructure the report with these sections:\n1. Object identification: concise neutral object name and one-sentence summary.\n2. Visual description: form, colors, pattern, components, and visible construction details.\n3. Condition observations: organize by structural fabric, beadwork, metal frame and clasp, surface soiling or staining, deformation, and previous repairs. For each observation, state its location precisely using the object\'s orientation as photographed.\n4. Condition priority table: assign each observed issue Low, Moderate, or High priority and briefly justify the rating.\n5. Handling, packing, and storage recommendations suitable for museum staff pending examination by a textile conservator.\n6. Recommended examination or documentation steps, clearly separating what can be concluded from the photograph from what requires in-person inspection.\n7. A 60–90 word collection-management summary.\n8. Concise accessibility alt text.\n\nUse professional conservation terminology, but label uncertain interpretations as tentative. Treat the scale and color checker only as reference aids; do not estimate exact dimensions or color values from the photograph.'
    }
  ]
})
import asyncio
import os

from runware import Runware


async def main():
    async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
        results = await client.run({
            "model": "anthropic:claude@opus-4.7",
            "settings": {
                "thinkingLevel": "adaptive"
            },
            "inputs": {
                "images": [
                    "https://assets.runware.ai/assets/inputs/019c3efd-2553-405e-ba95-0ac635f2f2bc.jpg"
                ]
            },
            "messages": [
                {
                    "role": "user",
                    "content": "Inspect the attached conservation photograph and draft a preliminary museum condition intake report for the evening bag. Ground every statement in visible evidence and do not invent provenance, dimensions, maker, date, materials, construction methods, or damage that cannot be supported by the image.\n\nStructure the report with these sections:\n1. Object identification: concise neutral object name and one-sentence summary.\n2. Visual description: form, colors, pattern, components, and visible construction details.\n3. Condition observations: organize by structural fabric, beadwork, metal frame and clasp, surface soiling or staining, deformation, and previous repairs. For each observation, state its location precisely using the object's orientation as photographed.\n4. Condition priority table: assign each observed issue Low, Moderate, or High priority and briefly justify the rating.\n5. Handling, packing, and storage recommendations suitable for museum staff pending examination by a textile conservator.\n6. Recommended examination or documentation steps, clearly separating what can be concluded from the photograph from what requires in-person inspection.\n7. A 60–90 word collection-management summary.\n8. Concise accessibility alt text.\n\nUse professional conservation terminology, but label uncertain interpretations as tentative. Treat the scale and color checker only as reference aids; do not estimate exact dimensions or color values from the photograph."
                }
            ]
        })


asyncio.run(main())
curl https://api.runware.ai/v1 \
  -H "Authorization: Bearer $RUNWARE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '[
    {
      "taskType": "textInference",
      "taskUUID": "a774184f-6ab1-4954-b3e9-82a9ab6c9ad2",
      "model": "anthropic:claude@opus-4.7",
      "settings": {
        "thinkingLevel": "adaptive"
      },
      "inputs": {
        "images": [
          "https://assets.runware.ai/assets/inputs/019c3efd-2553-405e-ba95-0ac635f2f2bc.jpg"
        ]
      },
      "messages": [
        {
          "role": "user",
          "content": "Inspect the attached conservation photograph and draft a preliminary museum condition intake report for the evening bag. Ground every statement in visible evidence and do not invent provenance, dimensions, maker, date, materials, construction methods, or damage that cannot be supported by the image.\n\nStructure the report with these sections:\n1. Object identification: concise neutral object name and one-sentence summary.\n2. Visual description: form, colors, pattern, components, and visible construction details.\n3. Condition observations: organize by structural fabric, beadwork, metal frame and clasp, surface soiling or staining, deformation, and previous repairs. For each observation, state its location precisely using the object's orientation as photographed.\n4. Condition priority table: assign each observed issue Low, Moderate, or High priority and briefly justify the rating.\n5. Handling, packing, and storage recommendations suitable for museum staff pending examination by a textile conservator.\n6. Recommended examination or documentation steps, clearly separating what can be concluded from the photograph from what requires in-person inspection.\n7. A 60–90 word collection-management summary.\n8. Concise accessibility alt text.\n\nUse professional conservation terminology, but label uncertain interpretations as tentative. Treat the scale and color checker only as reference aids; do not estimate exact dimensions or color values from the photograph."
        }
      ]
    }
  ]'
runware run anthropic:claude@opus-4.7 \
  settings.thinkingLevel=adaptive \
  inputs.images.0=https://assets.runware.ai/assets/inputs/019c3efd-2553-405e-ba95-0ac635f2f2bc.jpg \
  messages.0.role=user \
  messages.0.content="Inspect the attached conservation photograph and draft a preliminary museum condition intake report for the evening bag. Ground every statement in visible evidence and do not invent provenance, dimensions, maker, date, materials, construction methods, or damage that cannot be supported by the image.

Structure the report with these sections:
1. Object identification: concise neutral object name and one-sentence summary.
2. Visual description: form, colors, pattern, components, and visible construction details.
3. Condition observations: organize by structural fabric, beadwork, metal frame and clasp, surface soiling or staining, deformation, and previous repairs. For each observation, state its location precisely using the object's orientation as photographed.
4. Condition priority table: assign each observed issue Low, Moderate, or High priority and briefly justify the rating.
5. Handling, packing, and storage recommendations suitable for museum staff pending examination by a textile conservator.
6. Recommended examination or documentation steps, clearly separating what can be concluded from the photograph from what requires in-person inspection.
7. A 60–90 word collection-management summary.
8. Concise accessibility alt text.

Use professional conservation terminology, but label uncertain interpretations as tentative. Treat the scale and color checker only as reference aids; do not estimate exact dimensions or color values from the photograph."
{
  "taskType": "textInference",
  "taskUUID": "a774184f-6ab1-4954-b3e9-82a9ab6c9ad2",
  "model": "anthropic:claude@opus-4.7",
  "settings": {
    "thinkingLevel": "adaptive"
  },
  "inputs": {
    "images": [
      "https://assets.runware.ai/assets/inputs/019c3efd-2553-405e-ba95-0ac635f2f2bc.jpg"
    ]
  },
  "messages": [
    {
      "role": "user",
      "content": "Inspect the attached conservation photograph and draft a preliminary museum condition intake report for the evening bag. Ground every statement in visible evidence and do not invent provenance, dimensions, maker, date, materials, construction methods, or damage that cannot be supported by the image.\n\nStructure the report with these sections:\n1. Object identification: concise neutral object name and one-sentence summary.\n2. Visual description: form, colors, pattern, components, and visible construction details.\n3. Condition observations: organize by structural fabric, beadwork, metal frame and clasp, surface soiling or staining, deformation, and previous repairs. For each observation, state its location precisely using the object's orientation as photographed.\n4. Condition priority table: assign each observed issue Low, Moderate, or High priority and briefly justify the rating.\n5. Handling, packing, and storage recommendations suitable for museum staff pending examination by a textile conservator.\n6. Recommended examination or documentation steps, clearly separating what can be concluded from the photograph from what requires in-person inspection.\n7. A 60–90 word collection-management summary.\n8. Concise accessibility alt text.\n\nUse professional conservation terminology, but label uncertain interpretations as tentative. Treat the scale and color checker only as reference aids; do not estimate exact dimensions or color values from the photograph."
    }
  ]
}
Response
{
  "taskType": "textInference",
  "taskUUID": "a774184f-6ab1-4954-b3e9-82a9ab6c9ad2",
  "text": "# Preliminary Condition Intake Report\n\n## 1. Object Identification\n**Object name:** Beaded frame purse (evening bag / reticule with kiss-lock closure).\n**Summary:** A small rounded beaded bag with a metal frame and ball-clasp closure, featuring an all-over geometric beaded pattern in cream, brown, and dark red/purple tones, exhibiting multiple areas of loss, abrasion, and structural instability.\n\n## 2. Visual Description\n- **Form:** Rounded, roughly semi-circular to shield-shaped body suspended from a rigid metal top frame. The frame terminates in a central twin-ball kiss-lock clasp.\n- **Frame:** Metal (tentatively a white metal or silvered alloy based on tone), with visible decorative stamped or chased ornament along the upper band and hinged pivot points at each end.\n- **Beadwork:** Densely worked seed beads covering the entire visible face. The pattern comprises concentric lozenge/diamond motifs arranged in a lattice, alternating cream-white centers framed by golden-brown/amber and dark red-purple bead courses. A scalloped/undulating band of dark beads outlines each motif.\n- **Additional components:** A circular tan-colored applied element (tentatively leather or hide) is present in the lower-right quadrant, from which a short frayed fibrous tuft/tassel remnant extends. Loose fibrous strands are also visible along the lower edge.\n\n## 3. Condition Observations\n\n**Structural fabric / ground textile**\n- Exposed tan ground material visible at the upper-right shoulder of the bag where beadwork is lost, indicating localized separation of beads from substrate.\n- Additional smaller exposed area at the upper-left curve of the frame juncture.\n- Loose fibers protruding along the lower edge and lower-right, suggesting weakened seam threads or a degraded tassel/fringe attachment.\n\n**Beadwork**\n- Bead loss concentrated at the upper-right shoulder (largest loss area), with smaller losses at the upper-left near the frame.\n- Scattered thread slackening tentatively indicated by slight irregularity in bead alignment along the lower-left curve.\n- Overall bead surface appears intact in color and finish across the central field; no obvious large-scale iridescence loss discernible from the image.\n\n**Metal frame and clasp**\n- Frame surface shows uneven tonality and possible tarnish or patina, most visible along the upper edge and around the clasp.\n- A darker discoloration/possible residue is visible on the inner face of the frame just left of the clasp (tentative: tarnish, adhesive residue, or accreted soil).\n- Kiss-lock ball clasp appears present and aligned; functional status cannot be confirmed from the photograph.\n- Hinge points at both frame termini are visible; integrity not verifiable from image alone.\n\n**Surface soiling / staining**\n- General overall dulling consistent with age-related soiling.\n- The tan circular applied element in the lower-right shows tonal variation that may indicate staining, wear, or original patina (tentative).\n\n**Deformation**\n- Slight asymmetry between the left and right lower curves; whether this reflects original construction or deformation of the frame/textile is not determinable from a single view.\n- No pronounced creasing or collapse evident in this view.\n\n**Previous repairs**\n- No unambiguous evidence of prior repair visible; the tan circular applied patch in the lower-right may represent either an original decorative element (e.g., tassel anchor) or a later intervention — this is **tentative** and requires in-person assessment.\n\n## 4. Condition Priority Table\n\n| # | Observation | Location | Priority | Justification |\n|---|---|---|---|---|\n| 1 | Bead loss with exposed substrate | Upper-right shoulder | **High** | Active loss zone; adjacent beads at risk of cascading detachment. |\n| 2 | Secondary bead loss | Upper-left near frame | Moderate | Smaller area but in a stress zone adjacent to the rigid frame. |\n| 3 | Loose/frayed fibers and tuft remnant | Lower-right and lower edge | Moderate | Indicates degraded thread or fringe attachment; vulnerable to snagging. |\n| 4 | Possible thread slackening in beadwork | Lower-left curve | Moderate | Precursor to further bead loss if handled. |\n| 5 | Metal frame tarnish / discoloration | Upper frame and area beside clasp | Low | Cosmetic; stable if kept in controlled RH. |\n| 6 | Uncertain status of clasp/hinge function | Frame termini and central clasp | Moderate | Operability unknown; forcing risks damage. |\n| 7 | Overall surface soiling | Entire face | Low | Non-active; addressed only after conservator assessment. |\n| 8 | Possible minor asymmetry/deformation | Lower body outline | Low | Not confirmed; monitor. |\n\n## 5. Handling, Packing, and Storage Recommendations\n- Handle with clean nitrile gloves; support the object fully from beneath on a padded board — never lift by the frame alone, as the beaded body's weight may stress weakened attachment threads.\n- Do **not** open or operate the kiss-lock clasp until examined by a conservator.\n- Store flat, face-up, on an inert padded support (e.g., washed cotton stockinette over polyester batting on acid-free board), with a shaped ethafoam or rolled tissue support beneath the body to maintain its three-dimensional profile without pressure on the beadwork.\n- Cover loosely with acid-free tissue; avoid direct contact between tissue and loose bead areas — use a raised tissue tent if needed.\n- House in an acid-free box slightly larger than the object; label the exterior to indicate fragile beadwork and loose components.\n- Maintain stable environment: ~18–20 °C, 45–55% RH, low light (<50 lux when on view); isolate from sulfur-containing materials to limit further metal tarnishing.\n- Collect and retain any detached beads or fibers in a labeled archival vial or polyethylene bag stored with the object.\n\n## 6. Recommended Examination and Documentation Steps\n\n**Concluded from the photograph:**\n- Presence and approximate location of bead loss zones.\n- General pattern, palette, and construction type (beaded body on textile ground, metal kiss-lock frame).\n- Presence of loose fibers and an applied circular element in the lower-right.\n- Evidence of surface soiling and possible frame tarnish.\n\n**Requires in-person inspection:**\n- Verification of bead type, ground textile fiber, thread type, and lining materials.\n- Frame metal identification (visual, and where appropriate, non-destructive analytical testing).\n- Assessment of clasp and hinge mechanical function.\n- Interior examination for lining condition, maker's marks, labels, or inscriptions.\n- Reverse-side condition (not visible in this image).\n- Confirmation of whether the tan circular element is original or a repair.\n- Dimensions, weight, and detailed photographic documentation (raking light, macro of loss areas, UV examination for adhesives/repairs).\n- Stitch/thread analysis under magnification to assess extent of thread degradation.\n\n## 7. Collection-Management Summary\nThis beaded frame purse presents in fair but fragile condition. The beaded lattice pattern and metal kiss-lock frame remain visually coherent, but localized bead loss at the upper-right shoulder, secondary loss near the upper-left frame, and loose fibers along the lower edge indicate active vulnerability. Priority actions are stabilized housing, restricted handling, and referral to a textile conservator for full examination, substrate and thread assessment, and treatment planning. Provenance, date, materials, and maker are undetermined and require further research.\n\n## 8. Alt Text\nA small rounded beaded evening bag with a metal kiss-lock frame, photographed flat against a grey background beside a scale and color reference card. The bag is decorated with a geometric lattice of cream, amber-brown, and dark red-purple seed beads, and shows areas of bead loss at the upper right and loose fibers along the lower edge.",
  "cost": 0.078035,
  "finishReason": "stop"
}

text-generation

Payment Webhook Reliability Remediation Plan$0.51~3m 40s
# Northstar Billing — Payment Integrity Remediation Design

**Status:** Proposed, pending sign-off (Engineering, SRE, Security, FinOps, Support)
**Author:** Principal Engineer, Payments Platform
**Version:** 1.0

---

## 0. Terminology and Provenance

**Delivery semantics used throughout:**
- **Transport delivery**: at-least-once. Explicitly not deduplicated by SNS/SQS.
- **Ingress persistence**: one durable journal row per received delivery, including duplicates.
- **Canonical event materialization**: exactly one canonical row per idempotency key (enforced by unique index in PostgreSQL).
- **Business-effect execution**: **effectively-once** — each business effect (ledger posting, entitlement mutation, email send) is applied at most once per canonical event via a per-effect idempotency record co-committed with the effect (PostgreSQL) or via a conditional write keyed on the same identifier (DynamoDB).

We do **not** claim exactly-once end-to-end. We claim: (a) at-least-once transport, (b) exactly-once canonicalization, (c) effectively-once business effects.

**Assumptions introduced (not supplied in brief), flagged where used:**
- **A1**: Adyen `pspReference` + `eventCode` is stable per merchant account; PayPal `id` (event resource) is stable per merchant; Stripe `evt_...` is stable per Stripe account. Requires verification during Phase 0.
- **A2**: Merchant account identifier can be derived from webhook path or payload (`account` header for Stripe, `merchantAccountCode` for Adyen, `resource.supplementary_data.related_ids` for PayPal). Ingress must be extended to extract it. Verified per provider.
- **A3**: Payload compressed size (gzip) averages 2 KB, max 32 KB. Requires measurement.
- **A4**: RDS instance is db.r6g.4xlarge or larger with provisioned IOPS. If not, sizing in §8 changes.
- **A5**: Existing signature validation cannot be bypassed to accept a delivery; a signature failure produces a journal row with `signature_valid=false` and no canonical event.

---

## 1. Executive Summary and Success Criteria

### 1.1 Summary

The current pipeline treats webhook receipt, deduplication, and business-effect execution as loosely coupled operations with no shared idempotency contract. Ledger's "SELECT then INSERT" duplicate check is a classic time-of-check/time-of-use race; DynamoDB updates in Entitlements have no idempotency guard at all; and the internal event lacks merchant account scope, causing legitimate events to collide.

The remediation introduces:

1. **A durable delivery journal** in PostgreSQL, written synchronously before ACK-ing the provider (source of truth for "we received it").
2. **A canonical event table** with a **collision-safe idempotency key** unique on `(provider, merchant_account_id, provider_event_id_or_hash)`, materialized once per event.
3. **Per-consumer business-effect idempotency records** co-committed with the effect in PostgreSQL, and **conditional writes** keyed by the same effect ID in DynamoDB.
4. **A payment-lifecycle state machine** in the Ledger that distinguishes transport duplication (dropped) from out-of-order state transitions (buffered and reconciled, never silently discarded).
5. **A migration path** that dual-writes and shadow-validates without pausing intake.

### 1.2 Success Criteria (measurable)

| # | Criterion | Target | Measurement |
|---|-----------|--------|-------------|
| S1 | Duplicate business effects per 10⁶ deliveries | 0 (bounded by DB uniqueness) | Monthly reconciliation report |
| S2 | Legitimate events dropped due to key collision | 0 | Reconciliation vs. provider dashboards |
| S3 | Added ingress latency p99 | ≤ 80 ms | Ingress histogram |
| S4 | Planned write downtime during migration | ≤ 5 min | Change record |
| S5 | Journal completeness (deliveries recorded / deliveries acknowledged) | 100.000% | Journal count vs. ALB access logs |
| S6 | Time to replay a single event, controlled | ≤ 60 s from operator command | Runbook drill |
| S7 | Reconciliation report generated per tenant per day | ≥ 99% success | Batch job SLO |

---

## 2. Failure Analysis

### 2.1 The race condition in Ledger

```
T1: SELECT 1 FROM ledger WHERE provider_event_id = 'evt_X'  -> 0 rows
T2: SELECT 1 FROM ledger WHERE provider_event_id = 'evt_X'  -> 0 rows
T1: INSERT ... 'evt_X'  -> OK
T2: INSERT ... 'evt_X'  -> OK   ← duplicate
```
No unique constraint, no `SELECT ... FOR UPDATE`, no `INSERT ... ON CONFLICT`. Under SNS retry bursts, concurrent consumers observe the same pre-insert state. This produced the 286 duplicate ledger entries.

### 2.2 Collision domain error

Provider event IDs are unique **within a merchant account**, not globally. Ledger's dedup key was `provider_event_id` alone. When two merchant accounts emitted `evt_1a2b3c` for legitimate distinct events, the second was suppressed as a "duplicate." Root cause: `merchant_account_id` was never in the internal event.

### 2.3 Retry behavior

SNS standard + SQS standard: at-least-once with unbounded duplicate window in practice (SQS visibility timeouts, DLQ replays, manual replay up to 7 days). During the 47-minute degradation, SNS retried aggressively; each retry was a new SQS message. There is no message-ID-level dedup across retries.

### 2.4 Out-of-order events

Payment lifecycle (`authorized → captured → refunded`) can be observed in any order because: (a) providers emit events on independent workers, (b) SNS/SQS standard reorder freely, (c) retries of an early event may arrive after a later event succeeded. A naive "apply in arrival order" ledger accepts a refund against a payment that has no capture record, then rejects the capture as invalid — silently losing state.

### 2.5 Why SNS/SQS FIFO alone is insufficient

FIFO with content-based dedup would fix **transport duplicates within a 5-minute window** for a single `MessageGroupId`. It does not solve:

1. **7-day manual replay** exceeds FIFO's 5-minute dedup window.
2. **Consumer-side idempotency**: FIFO delivery is still at-least-once (visibility timeout expiry re-delivers). Business-effect dedup still required.
3. **Ordering scope**: FIFO orders within a `MessageGroupId`. To preserve payment ordering, group must be `payment_id`, but that serializes all events per payment and destroys the fan-out throughput (10k/s burst → thousands of groups is fine, but a hot payment blocks). Also SNS→SQS FIFO does not fan out to standard queues, forcing every consumer to FIFO.
4. **Cross-consumer idempotency**: FIFO dedup is per-queue. Ledger and Entitlements would each need independent guarantees.
5. **Collision domain**: FIFO dedup uses message content hash. It does not know that provider event IDs are only unique per merchant.
6. **Journal requirement**: FinOps needs *every* delivery recorded, including duplicates. FIFO discarding duplicates before the journal would violate this.

**Decision**: keep SNS/SQS standard for fan-out; solve idempotency and ordering in PostgreSQL where the constraints are enforceable and auditable.

---

## 3. Target Architecture

### 3.1 Data flow

```
Provider → ALB → Ingress API (Go)
                    │
                    ├─(1) verify signature
                    ├─(2) extract merchant_account_id, build envelope
                    ├─(3) BEGIN tx
                    │      INSERT delivery_journal (always)
                    │      INSERT canonical_event  ON CONFLICT DO NOTHING
                    │    COMMIT
                    ├─(4) if canonical row was new → publish to SNS
                    └─(5) return HTTP 200

SNS topic (standard, fan-out)
  ├→ SQS: ledger.q         → Ledger workers → PostgreSQL (ledger + effect_idem in same tx)
  ├→ SQS: entitlements.q   → Entitlements workers → DynamoDB (conditional write) + effect_idem in Postgres
  ├→ SQS: email.q          → (unchanged in phase 1, tolerates duplicates)
  └→ SQS: analytics.q      → (unchanged in phase 1)

Async:
  - Journal archiver → S3 (Parquet, partitioned by date) after 90 days
  - Reconciliation job → S3 report per tenant per day
  - Replay CLI → publishes canonical events back onto SNS with replay flag
```

### 3.2 Transactional boundaries and sources of truth

| Concern | Source of truth | Transactional mechanism |
|---|---|---|
| "Did we receive this delivery?" | `delivery_journal` (Postgres) | Single-row insert, committed before HTTP 200 |
| "Have we canonicalized this event?" | `canonical_event` (Postgres) | Unique index on idempotency key |
| "Has this ledger effect been applied?" | `ledger_effect_idem` + `ledger_entry` (Postgres) | Same transaction |
| "Has this entitlement effect been applied?" | `entitlement_effect_idem` (Postgres) is authoritative; DynamoDB is projection | Postgres commit first, then DynamoDB conditional write; recovery loop reconciles |
| "Current payment lifecycle state" | `payment_state` (Postgres) | Row-level lock + state-transition validation |
| Cache | Redis | Never source of truth; may be flushed |

### 3.3 Why the journal-before-ACK trade-off is acceptable

Adding a single Postgres insert to the ingress path costs ~3–8 ms p50, ~20–40 ms p99 on a well-sized RDS with connection pooling (PgBouncer transaction mode). Well within the 80 ms budget. The alternative — journal asynchronously — risks losing the audit record if the API crashes between ACK and journal, which FinOps has ruled out.

### 3.4 Publish-after-commit

SNS publish happens **after** the Postgres commit. If SNS publish fails, the canonical row exists but no message is emitted. A janitor process scans `canonical_event WHERE published_at IS NULL AND created_at < now() - interval '30 seconds'` and re-publishes. This is safe because consumers are idempotent on canonical event ID.

---

## 4. Canonical Event Envelope

### 4.1 Envelope (JSON, versioned)

| Field | Type | Req | Semantics |
|---|---|---|---|
| `schema_version` | string (semver) | Y | Envelope schema, e.g. `"1.0"` |
| `event_id` | UUIDv7 | Y | Northstar-assigned; primary key of `canonical_event` |
| `idempotency_key` | string (≤ 200 chars) | Y | See §4.2 |
| `provider` | enum(`stripe`,`adyen`,`paypal`) | Y | |
| `merchant_account_id` | string | Y | Provider-scoped account (Stripe account, Adyen merchant, PayPal payer) |
| `tenant_id` | UUID | Y | Northstar tenant |
| `provider_event_id` | string | N | Provider's ID if supplied; null if provider omits |
| `event_type` | string | Y | Canonicalized, e.g. `payment.captured` |
| `payment_id` | string | N | Provider payment/charge ID; required for lifecycle events |
| `amount_minor` | int64 | N | Minor units (cents); required for money-moving events |
| `currency` | ISO-4217 string | N | Required with `amount_minor` |
| `provider_event_time` | RFC3339 | Y | From provider |
| `received_at` | RFC3339 | Y | Ingress receipt time (server clock) |
| `payload_ref` | string | Y | S3 URI + SHA-256 of minimized payload; see §12 |
| `payload_digest` | hex(SHA-256) | Y | Hash of full raw payload (for audit, without storing it) |
| `signature_valid` | bool | Y | Result of provider signature check |
| `producer_version` | string | Y | Ingress build hash |

`payload` is **not** in the envelope on the wire to downstream consumers; they fetch minimized fields via `payload_ref` if needed. Reduces SNS message size, PII exposure, and cost.

### 4.2 Idempotency key

```
idempotency_key = provider + ":" + merchant_account_id + ":" + event_scope

where event_scope =
  if provider_event_id present:  provider_event_id
  else:                          "syn:" + sha256_hex(
                                    payment_id || event_type ||
                                    provider_event_time_iso ||
                                    amount_minor || currency ||
                                    canonical_provider_specific_fields
                                 )
```

- **Collision safety**: including `merchant_account_id` fixes the observed suppression bug.
- **Missing provider event ID (synthesized keys)**: PayPal `TRANSACTION.SALE.COMPLETED` under some flows lacks a stable event ID; synthesize from stable business fields. `syn:` prefix makes provenance explicit and reconciliable.
- **Length**: ≤ 200 chars fits Stripe (`evt_` + 24), Adyen (`pspReference` up to 16 + eventCode), PayPal IDs, plus prefix.
- **Stability**: derived only from provider-supplied stable fields; never includes our timestamps or UUIDs.

### 4.3 Handling providers with no stable ID

Synthesized keys are inherently weaker: if the provider re-sends the same event with the same fields, we correctly dedupe; if they mutate a field on retry (rare, but Adyen has been observed to), we would double-count. Mitigation:

1. Per-provider allowlist of fields used in synthesis, reviewed quarterly.
2. Reconciliation report flags synthetic-key events with retry counts > 1 for manual review.
3. Prefer provider event ID everywhere it exists.

---

## 5. PostgreSQL Schema

### 5.1 Delivery journal (append-only, immutable)

```sql
CREATE TABLE delivery_journal (
    delivery_id         UUID        NOT NULL DEFAULT gen_random_uuid(),
    received_at         TIMESTAMPTZ NOT NULL,
    provider            TEXT        NOT NULL,
    merchant_account_id TEXT,
    provider_event_id   TEXT,
    tenant_id           UUID,
    signature_valid     BOOLEAN     NOT NULL,
    payload_digest      BYTEA       NOT NULL,          -- sha256 of raw body
    payload_bytes       INT         NOT NULL,
    source_ip           INET,
    http_headers_digest BYTEA,
    idempotency_key     TEXT,                          -- null if signature invalid or unparseable
    canonical_event_id  UUID,                          -- null if not canonicalized
    outcome             TEXT NOT NULL,                 -- 'accepted_new','accepted_duplicate','rejected_signature','rejected_parse'
    PRIMARY KEY (received_at, delivery_id)
) PARTITION BY RANGE (received_at);

-- Monthly partitions; pg_partman manages creation and detach.
CREATE INDEX ON delivery_journal (idempotency_key) WHERE idempotency_key IS NOT NULL;
CREATE INDEX ON delivery_journal (canonical_event_id) WHERE canonical_event_id IS NOT NULL;
CREATE INDEX ON delivery_journal (provider, received_at);
```

**Retention**: Operational metadata may be archived after 90 days (per brief). Journal partitions older than 90 days are exported to S3 (Parquet, KMS-encrypted, Object Lock in compliance mode for 7 years) then detached and dropped. The S3 export **is** the 7-year record for delivery metadata. Money-moving records live in `ledger_entry` (7-year retention in Postgres + logical backups).

### 5.2 Canonical events

```sql
CREATE TABLE canonical_event (
    event_id            UUID        PRIMARY KEY,       -- UUIDv7
    idempotency_key     TEXT        NOT NULL,
    provider            TEXT        NOT NULL,
    merchant_account_id TEXT        NOT NULL,
    tenant_id           UUID        NOT NULL,
    event_type          TEXT        NOT NULL,
    payment_id          TEXT,
    amount_minor        BIGINT,
    currency            CHAR(3),
    provider_event_time TIMESTAMPTZ NOT NULL,
    received_at         TIMESTAMPTZ NOT NULL,
    payload_ref         TEXT        NOT NULL,
    schema_version      TEXT        NOT NULL,
    published_at        TIMESTAMPTZ,
    created_at          TIMESTAMPTZ NOT NULL DEFAULT now()
);

CREATE UNIQUE INDEX canonical_event_idem
    ON canonical_event (idempotency_key);
CREATE INDEX canonical_event_payment
    ON canonical_event (tenant_id, payment_id)
    WHERE payment_id IS NOT NULL;
CREATE INDEX canonical_event_publish_pending
    ON canonical_event (created_at)
    WHERE published_at IS NULL;
```

Not partitioned initially; sized in §8. Add range partitioning by `created_at` when row count exceeds ~500M (est. ~3 years).

### 5.3 Processing attempts

```sql
CREATE TABLE processing_attempt (
    attempt_id      BIGSERIAL PRIMARY KEY,
    event_id        UUID        NOT NULL REFERENCES canonical_event(event_id),
    consumer        TEXT        NOT NULL,            -- 'ledger','entitlements',...
    worker_id       TEXT        NOT NULL,
    started_at      TIMESTAMPTZ NOT NULL DEFAULT now(),
    finished_at     TIMESTAMPTZ,
    outcome         TEXT,                            -- 'applied','skipped_duplicate','deferred','failed'
    error_class     TEXT,
    error_message   TEXT
);
CREATE INDEX ON processing_attempt (event_id, consumer);
CREATE INDEX ON processing_attempt (consumer, started_at)
    WHERE finished_at IS NULL;                       -- stuck-work detection
```

### 5.4 Per-consumer business-effect idempotency

```sql
CREATE TABLE ledger_effect_idem (
    effect_id       TEXT        PRIMARY KEY,          -- see §6.1
    event_id        UUID        NOT NULL REFERENCES canonical_event(event_id),
    ledger_entry_id UUID        NOT NULL,             -- FK to ledger_entry
    applied_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);

CREATE TABLE entitlement_effect_idem (
    effect_id       TEXT        PRIMARY KEY,
    event_id        UUID        NOT NULL REFERENCES canonical_event(event_id),
    dynamo_pk       TEXT        NOT NULL,
    dynamo_sk       TEXT        NOT NULL,
    intent          JSONB       NOT NULL,             -- what we intended to apply
    ddb_confirmed_at TIMESTAMPTZ,                     -- set after DynamoDB confirms
    applied_at      TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX ON entitlement_effect_idem (ddb_confirmed_at)
    WHERE ddb_confirmed_at IS NULL;                   -- recovery worker scans
```

### 5.5 Payment lifecycle state

```sql
CREATE TABLE payment_state (
    tenant_id           UUID  NOT NULL,
    provider            TEXT  NOT NULL,
    merchant_account_id TEXT  NOT NULL,
    payment_id          TEXT  NOT NULL,
    state               TEXT  NOT NULL,               -- 'unknown','authorized','captured','refunded','voided','failed'
    last_event_time     TIMESTAMPTZ NOT NULL,
    pending_events      JSONB NOT NULL DEFAULT '[]',  -- buffered out-of-order events
    updated_at          TIMESTAMPTZ NOT NULL DEFAULT now(),
    PRIMARY KEY (tenant_id, provider, merchant_account_id, payment_id)
);
```

---

## 6. Consumer Transaction Logic

### 6.1 Effect ID construction

```
effect_id = consumer_name + ":" + event_id + ":" + effect_kind
```
E.g. `ledger:018f...:posting`, `entitlements:018f...:extend`. Deterministic from canonical event; safe as primary key.

### 6.2 Ledger consumer pseudocode

```
on_message(msg):
    env = parse(msg)
    tx = pg.begin(SERIALIZABLE)          # or READ COMMITTED with ON CONFLICT; see below
    try:
        # 1. Claim the work with a processing_attempt row
        attempt_id = tx.insert(processing_attempt, {
            event_id: env.event_id, consumer:'ledger',
            worker_id: WORKER_ID, started_at: now()
        })

        # 2. Advisory lock scoped to payment to serialize state transitions
        tx.exec("SELECT pg_advisory_xact_lock(hashtextextended($1,0))",
                env.tenant_id || env.payment_id)

        # 3. Load and validate lifecycle transition
        state = tx.select_for_update(payment_state, keys...)
        decision = lifecycle.evaluate(state, env)     # apply | buffer | ignore_duplicate
        if decision == ignore_duplicate:
            tx.update(processing_attempt, id=attempt_id,
                      outcome='skipped_duplicate', finished_at=now())
            tx.commit(); ack(msg); return

        if decision == buffer:
            state.pending_events = append(state.pending_events, env.minified())
            tx.update(payment_state, state)
            tx.update(processing_attempt, id=attempt_id,
                      outcome='deferred', finished_at=now())
            tx.commit(); ack(msg); return            # will be re-drained when predecessor arrives

        # decision == apply: also drain any pending events now applicable
        events_to_apply = [env] + lifecycle.drain(state, env)

        for e in events_to_apply:
            effect_id = "ledger:" + e.event_id + ":posting"
            # 4. Atomic effect: idem-row + ledger row in one tx
            r = tx.exec("""
                INSERT INTO ledger_effect_idem(effect_id, event_id, ledger_entry_id)
                VALUES ($1, $2, $3)
                ON CONFLICT (effect_id) DO NOTHING
                RETURNING effect_id
            """, effect_id, e.event_id, new_uuid)
            if r.rows == 0:
                continue                              # already applied
            tx.insert(ledger_entry, ...)              # references ledger_entry_id above
            state = lifecycle.advance(state, e)

        tx.update(payment_state, state)
        tx.update(processing_attempt, id=attempt_id,
                  outcome='applied', finished_at=now())
        tx.commit()
        ack(msg)
    except SerializationFailure:
        tx.rollback(); nack(msg)                      # SQS visibility timeout will redeliver
    except Exception as e:
        tx.rollback()
        record_failed_attempt(env.event_id, e)
        nack(msg)                                     # DLQ after N attempts
```

**Concurrency correctness**: `pg_advisory_xact_lock` on `payment_id` serializes per-payment work. `INSERT ... ON CONFLICT DO NOTHING` on `ledger_effect_idem` (primary key) is atomic even without the advisory lock; the lock exists to serialize state-machine transitions, not idempotency.

**Worker death recovery**: An orphaned `processing_attempt` with `finished_at IS NULL` older than 10 minutes is swept by a reaper that (a) checks whether the effect row exists — if yes, marks attempt `applied`; if no, marks `failed` and re-enqueues by publishing to the retry SNS topic. Since the effect row is the source of truth, the reaper cannot cause duplicates.

### 6.3 Entitlements consumer pseudocode

DynamoDB cannot participate in the Postgres transaction. Pattern: **outbox in Postgres, then apply, then confirm.**

```
on_message(msg):
    env = parse(msg)
    effect_id = "entitlements:" + env.event_id + ":extend"
    intent = build_intent(env)                        # deterministic

    # Phase A: reserve the intent atomically in Postgres
    tx = pg.begin()
    r = tx.exec("""
        INSERT INTO entitlement_effect_idem(effect_id, event_id, dynamo_pk, dynamo_sk, intent)
        VALUES ($1, $2, $3, $4, $5)
        ON CONFLICT (effect_id) DO NOTHING
        RETURNING effect_id, ddb_confirmed_at
    """, effect_id, env.event_id, intent.pk, intent.sk, intent)
    if r.rows == 0:
        row = tx.select_one(entitlement_effect_idem, effect_id=effect_id)
        tx.commit()
        if row.ddb_confirmed_at is not None:
            ack(msg); return                          # fully done
        # else: fall through to Phase B (retrying an unconfirmed prior attempt)
    else:
        tx.commit()

    # Phase B: apply to DynamoDB with a conditional write keyed on effect_id
    try:
        ddb.update_item(
            TableName='entitlements',
            Key={'pk': intent.pk, 'sk': intent.sk},
            UpdateExpression='ADD applied_effects :eid SET expires_at = :exp',
            ConditionExpression='NOT contains(applied_effects, :eid)',
            ExpressionAttributeValues={
                ':eid': effect_id,
                ':exp': intent.new_expiry,
            },
        )
    except ConditionalCheckFailedException:
        pass                                          # already applied to DDB previously

    # Phase C: confirm
    pg.exec("UPDATE entitlement_effect_idem SET ddb_confirmed_at = now() WHERE effect_id = $1",
            effect_id)
    ack(msg)
```

**Correctness**:
- If crash after Phase A, before Phase B: next redelivery finds `ddb_confirmed_at IS NULL`, retries Phase B. DynamoDB conditional write is idempotent on `effect_id`.
- If crash after Phase B, before Phase C: next redelivery retries Phase B; conditional check fails harmlessly; Phase C runs.
- If DynamoDB accepts write but Postgres update fails: effect exists in DDB, next redelivery sees unconfirmed idem row, re-attempts DDB (idempotent), confirms in Postgres.
- Recovery worker scans `entitlement_effect_idem WHERE ddb_confirmed_at IS NULL AND applied_at < now() - interval '5 minutes'` and drives Phase B+C without an inbound message.

The `applied_effects` set in DynamoDB has bounded growth per entitlement item; a TTL/GC job trims effect IDs older than 30 days once the corresponding Postgres row is confirmed.

---

## 7. Out-of-Order Lifecycle State Machine

### 7.1 Two concerns, separated

| Concern | Mechanism | Table |
|---|---|---|
| Transport duplicate | Same idempotency key already canonicalized | `canonical_event` unique index |
| Business duplicate effect | Same `effect_id` already applied | `*_effect_idem` |
| Out-of-order transition | State machine defers, buffers, and later drains | `payment_state.pending_events` |

Transport dedup **must not** silently discard out-of-order events. A refund arriving before its capture is not a duplicate — it is a valid future event.

### 7.2 State machine

States: `unknown → authorized → captured → refunded (terminal) | voided (terminal) | failed (terminal)`

Transition table (excerpt):

| From \ Event | authorized | captured | refunded | voided | failed |
|---|---|---|---|---|---|
| unknown | → authorized | **buffer** | **buffer** | **buffer** | → failed |
| authorized | ignore (dup) | → captured | **buffer** | → voided | → failed |
| captured | ignore | ignore | → refunded | reject-invalid | reject-invalid |
| refunded | ignore | ignore | ignore/partial-refund handling | reject | reject |

`buffer` = write into `payment_state.pending_events`; do not apply ledger effect yet; ACK message (canonical event is durable, effect will materialize when the predecessor arrives).

`ignore` = we've seen this event class already; treat as duplicate (but still record `processing_attempt` with `outcome='skipped_duplicate'`).

`reject-invalid` = state contradiction. Do **not** silently drop. Route to `payment_state_conflict` table for FinOps review, alert fires.

### 7.3 Buffering bound

`pending_events` capped at 20 entries per payment (way beyond expected 3–5). If exceeded, alert fires: indicates provider misbehavior or missing predecessor event that requires manual intervention. Bounded buffer prevents unbounded row growth.

### 7.4 Time-based drain

A background sweeper runs every 5 minutes: for `payment_state` rows with `pending_events` non-empty and `updated_at < now() - interval '1 hour'`, it emits a `payment_state_stalled` metric per (tenant, provider) and creates a Support ticket for the specific payment. Prevents silent losses when the predecessor event never arrives.

---

## 8. Capacity Reasoning

### 8.1 Volumes

- Sustained: 2,000 req/s → **172.8M deliveries/day**.
- Burst: 10,000 req/s × 5 min = 3M deliveries in 5 min.
- Assumed duplicate multiplier under normal ops: 1.05×. Under incident: up to 3× (matches evidence: 18,420 / 6,130 ≈ 3.0).

### 8.2 Journal write load

- Sustained: 2,000 inserts/s into `delivery_journal`. Each row ~250 bytes + 2 indexes ≈ 400 B written per row after WAL amplification.
- Burst: 10,000 inserts/s → ~4 MB/s write, ~40 MB/s WAL with fsync amplification. Well within db.r6g.4xlarge with gp3 12,000 IOPS.
- Monthly partition size at sustained: 172.8M × 30 × 250 B ≈ **1.3 TB/month** raw; ~2 TB with indexes. Partition + archive to S3 after 90 days keeps hot storage at ~6 TB.

**Requires load test**: sustained 10k/s insert into partitioned parent with 2 indexes on db.r6g.4xlarge. If p99 insert latency exceeds 30 ms, upgrade to r6g.8xlarge or shard journal by `provider` prefix (separate tables per provider, UNION view for reads).

### 8.3 Canonical event and effect writes

- Sustained ~1,900/s new canonical events (5% duplicates suppressed).
- Ledger effect insert rate: ~1,900/s. Effect table growth: 172.8M × 0.95 × ~150 B ≈ 25 GB/month. Manageable without partitioning for ~2 years.

### 8.4 Ingress latency budget

Path (target p99):

| Stage | Budget (ms) |
|---|---|
| TLS + ALB | 10 |
| Signature verify | 8 |
| Parse + envelope build | 5 |
| PgBouncer + tx (insert × 2, commit) | 35 |
| SNS publish (async, not on critical path if commit-then-publish outside request) | — |
| Slack | 22 |
| **Total** | **80** |

**Decision**: Return HTTP 200 after Postgres commit but **before** SNS publish. SNS publish happens on a background goroutine with retry, and the janitor process backstops failures. This keeps SNS latency (~15 ms p99, spike-prone) out of the provider-facing path.

### 8.5 Bottlenecks and mitigations

| Bottleneck | Symptom | Mitigation |
|---|---|---|
| RDS writer IOPS under burst | Journal insert p99 spike | Provisioned IOPS, PgBouncer transaction pool, batch inserts (100-row COPY per 20ms window in ingress) |
| Advisory-lock contention on hot payment | Ledger throughput drops for that payment | Acceptable; state transitions must serialize per payment |
| DynamoDB throttling on entitlement update | Phase B retries pile up | On-demand capacity; monitor `ThrottledRequests` |
| Postgres connection storm from EKS pods | `too many connections` | PgBouncer at ~1000 client → 200 server connections |

**Requires load test**: burst behavior, PgBouncer sizing, batched journal insert efficacy.

---

## 9. Migration Plan (Zero-Intake-Downtime)

### 9.1 Principles

- Never pause webhook intake.
- Dual-write during transition; canonical event is authoritative once shadow validation passes.
- Every phase has an explicit rollback that cannot produce duplicate effects.

### 9.2 Phases

**Phase 0 — Foundations (week 1–2)**
- Deploy new schema (journal, canonical_event, effect_idem, payment_state) to RDS with `CREATE TABLE IF NOT EXISTS`. No writes yet.
- Deploy PgBouncer.
- Extend ingress to extract `merchant_account_id` per provider; ship behind a feature flag with logging only.

**Phase 1 — Shadow journaling (week 3)**
- Ingress writes `delivery_journal` and `canonical_event` in shadow mode: rows written, SNS still uses the old envelope.
- Metric: `journal_write_error_rate`. Gate: < 0.01% for 72 hours.

**Phase 2 — Dual publish (week 4)**
- Ingress publishes both old envelope (unchanged topic) and new envelope (new SNS topic `payment-events-v2`).
- Ledger and Entitlements subscribe to v2 in **shadow mode**: they compute what they *would* do and log to `shadow_decision` table, without applying effects.
- Gate: reconcile shadow decisions vs. actual old-path decisions for 72 hours. Discrepancies triaged manually.

**Phase 3 — Ledger cutover (week 5)**
- Ledger v2 becomes authoritative; old ledger consumer stops. Feature-flagged per tenant, ramped 1% → 10% → 100% over 4 days.
- Effect idempotency prevents duplicate application even if a message is processed by both old and new consumers during ramp — this is the critical safety property.
- Rollback: disable v2 flag; old consumer resumes. Because effect_idem is checked by both (old consumer patched to read effect_idem on rollback), no duplicates.

**Phase 4 — Entitlements cutover (week 6)**
- Same pattern. DynamoDB conditional-write keyed on `effect_id` is the safety net.

**Phase 5 — Email & Analytics** (post-8-weeks; out of scope for first release)
- They continue consuming the v1 topic via a bridge that republishes canonical events in the old shape. When migrated, bridge is removed.

### 9.3 Backfill

Backfill of pre-migration data is **not** performed. Historical events remain in the legacy ledger. Canonical event table starts fresh at Phase 1. Reconciliation from Phase 1 forward.

### 9.4 Rollback criteria

Automatic rollback triggers:
- Ingress p99 latency > 100 ms for 10 minutes.
- `journal_write_error_rate` > 0.1% for 5 minutes.
- Canonical event uniqueness violations > 0.5% of inserts (indicates key derivation bug).
- Ledger v2 producing effects at > 1.5× v1 rate on same input stream (shadow discrepancy).

Rollback playbook: flip flag → verify old consumer resumes → confirm no gap in journal → post-mortem.

---

## 10. Observability

### 10.1 Metrics (Prometheus names)

| Metric | Type | Labels | Alert |
|---|---|---|---|
| `ingress_deliveries_total` | counter | provider, outcome | — |
| `ingress_latency_seconds` | histogram | provider | p99 > 80ms 5m |
| `journal_insert_errors_total` | counter | error_class | > 10/min |
| `canonical_event_duplicates_prevented_total` | counter | provider | Report only |
| `canonical_event_publish_pending` | gauge | — | > 1000 for 5m |
| `idempotency_key_collisions_suspected_total` | counter | provider | > 0 pages FinOps |
| `consumer_processed_total` | counter | consumer, outcome | — |
| `consumer_processing_seconds` | histogram | consumer | p99 > 2s 10m |
| `effect_apply_conflict_total` | counter | consumer | Rate change > 3σ |
| `payment_state_invalid_transition_total` | counter | provider, from_state, event | > 0 pages |
| `payment_state_pending_events_gauge` | gauge | tenant | > 5 per payment |
| `payment_state_stalled_total` | counter | provider | > 0/hour pages |
| `entitlement_ddb_unconfirmed_gauge` | gauge | — | > 100 for 10m |
| `processing_attempt_stuck_gauge` | gauge | consumer | > 50 for 10m |
| `replay_events_total` | counter | operator, reason | Report + audit log |
| `reconciliation_mismatch_total` | counter | tenant, category | > 0 opens ticket |

### 10.2 Structured log fields (mandatory on every log line in this pipeline)

`trace_id, span_id, delivery_id, event_id, idempotency_key, provider, merchant_account_id, tenant_id, payment_id (if any), consumer, outcome`

### 10.3 Traces

OpenTelemetry spans: `ingress.receive` → `ingress.verify` → `ingress.persist` → `ingress.publish` (async) → `consumer.claim` → `consumer.state_eval` → `consumer.apply` → `consumer.ack`. Propagate trace ID in envelope metadata.

### 10.4 Dashboards

- **Ingress health**: RPS, p50/p99 latency, journal error rate, signature failure rate.
- **Deduplication**: duplicates prevented per hour, ratio vs. total, per provider.
- **Consumer health**: throughput, effect apply rate, DLQ depth, stuck attempts.
- **Lifecycle**: state distribution, invalid transitions, pending events by age.
- **Reconciliation**: last successful run per tenant, open mismatches, aged mismatches.

### 10.5 Alerts (priority)

P1 pages: journal write errors > 0.1%, ingress p99 > 100ms sustained, `payment_state_invalid_transition_total` non-zero, `reconciliation_mismatch_total` for money-moving categories.
P2 tickets: DLQ depth > 100, unconfirmed DDB effects > 100, replay activity outside change window.

---

## 11. Reconciliation and Repair

### 11.1 Daily reconciliation job

Runs per tenant per day. Reads:
1. Provider records via API (Stripe `events.list`, Adyen report, PayPal reports) for date D.
2. `delivery_journal` for date D, filtered to that tenant.
3. `canonical_event` for date D.
4. `ledger_entry` for date D.

Produces a report with four sections:

| Section | Definition | Action |
|---|---|---|
| **A. Provider without journal** | Event exists in provider but not in journal | Investigate — likely missed webhook. May require manual replay. |
| **B. Journal without canonical** | Delivery accepted but not canonicalized | Investigate parse/signature bug. |
| **C. Canonical without ledger effect** (money-moving types) | Consumer failed to apply | Check DLQ, `processing_attempt`. |
| **D. Ledger effect without provider record** | Ghost entries | High severity; potential double-counting from pre-migration data. |

### 11.2 No automatic money movement

The job **never** posts corrections. It generates a signed report to S3, an entry in `reconciliation_report`, and if any section is non-empty, opens a Jira ticket assigned to FinOps. Corrections require:

1. FinOps analyst reviews report.
2. Two-person approval in the FinOps repair tool.
3. Repair emits an event through the same canonical pipeline with `event_type='ops.manual_correction'` and a distinct `effect_id` prefix. This appears in reconciliation on future runs, closing the loop.

### 11.3 Comparison keys

Reconciliation joins on `(provider, merchant_account_id, provider_event_id)` where available; falls back to `(payment_id, event_type, amount_minor, currency, provider_event_time)` for synthetic-key events. Discrepancies in the synthetic space are treated as suspected collisions and reviewed.

---

## 12. Security, Privacy, Retention

### 12.1 Payload minimization

Raw payloads are **not** stored in Postgres. On ingress:
1. Compute `payload_digest = sha256(raw_body)`; store in journal.
2. Extract only allowlisted fields per provider into a minimized JSON (`payment_id`, `amount`, `currency`, `event_type`, `merchant_account_id`, non-PII metadata).
3. Store minimized JSON in S3 (`s3://northstar-events/{tenant}/{yyyy/mm/dd}/{event_id}.json.gz`), SSE-KMS with per-tenant CMK.
4. `payload_ref` in canonical event is the S3 URI.

Full raw payload is retained transiently in the ingress process for signature verification, then discarded. No PII (card holder name, email, address) is persisted in Postgres.

### 12.2 Encryption

- RDS: encryption at rest (KMS), TLS in transit, client cert auth for consumer roles.
- S3: SSE-KMS with per-tenant CMK for minimized payloads; SSE-KMS with platform CMK for journal archive.
- DynamoDB: encryption at rest, VPC endpoints only.
- SNS/SQS: SSE-KMS.

### 12.3 Tenant isolation

- Row-level security on all tables keyed by `tenant_id`; consumer roles carry tenant context via `SET LOCAL app.tenant_id`.
- S3 prefix per tenant, IAM condition on prefix.
- DynamoDB partition key includes `tenant_id`.

### 12.4 Replay controls

Replay CLI requires:
- SSO with hardware key.
- Second reviewer approval (`replay-approver` group) for any replay > 1 event.
- Audit log to CloudTrail + `replay_audit` table with operator, reason, event IDs, ticket link.
- Hard limit: 10,000 events per replay command; larger replays require change ticket.
- Replays cannot bypass idempotency; effect_idem prevents double application.

### 12.5 Retention

| Data | Retention | Location | Mechanism |
|---|---|---|---|
| `ledger_entry` and related financial records | 7 years | Postgres + S3 backup | Retention policy documented, deletion requires FinOps + Legal |
| Delivery journal (hot) | 90 days | Postgres partitions | Auto-detach and archive |
| Delivery journal (archive) | 7 years | S3 Object Lock (compliance mode) | Immutable |
| Canonical event | 7 years | Postgres (partitioned after year 1) | |
| Minimized payloads | 7 years | S3 with lifecycle to Glacier at 90d | |
| Processing attempts | 90 days | Postgres | Truncated older partitions |
| Reconciliation reports | 7 years | S3 Object Lock | |

Deletions of PII per GDPR/CCPA operate on the minimized payload S3 objects and non-financial fields; financial records are retained under the legal-basis exemption and documented in the DPA.

---

## 13. Test Strategy

| Category | Test | Pass criteria |
|---|---|---|
| **Concurrency** | Fire 50 goroutines inserting same idempotency_key | Exactly 1 canonical row; 49 conflict errors handled |
| **Concurrency** | Fire 50 parallel Ledger consumers on same event | Exactly 1 `ledger_effect_idem` row, 1 `ledger_entry` |
| **Crash recovery** | Kill Ledger pod between insert and commit | Redelivery reprocesses; no duplicate; attempt reaper cleans stale row |
| **Crash recovery** | Kill Entitlements between Phase A and Phase B | DDB gets exactly one update, `ddb_confirmed_at` set |
| **Crash recovery** | Kill Entitlements between Phase B and Phase C | Idem row remains unconfirmed; recovery worker completes it; no duplicate DDB write |
| **Delayed duplicates** | Replay a 6-day-old event | Duplicate detected via idempotency_key; journal has 2 delivery rows, 1 canonical row |
| **Collision** | Two merchant accounts send `provider_event_id='evt_1'` for different payments | Both processed; distinct canonical rows |
| **Out-of-order** | Send refund before capture | Refund buffered; capture applied; then buffered refund drained; ledger has both entries in correct order |
| **Out-of-order** | Send refund with no capture ever arriving | After 1 hour, `payment_state_stalled` fires; ticket opened; refund not silently lost |
| **Partial failure** | DynamoDB throws for 10 min | Retries with backoff; recovery worker drives to completion; no duplicate |
| **Replay** | Operator replays 1,000 events already applied | 0 new effects; `replay_audit` populated; metrics show duplicates prevented |
| **Migration compat** | Old and new consumers both active on same event | Exactly 1 effect (proven by effect_idem constraint) |
| **Burst load** | 10k/s for 5 min, 10% synthetic duplicates | Ingress p99 < 80 ms; 0 message loss; journal complete |
| **Signature failure** | Send invalid signature | Journal row with `signature_valid=false`, no canonical event, no downstream propagation |
| **Fuzz** | Malformed payloads, oversized bodies, truncated JSON | Ingress rejects with 4xx, journal row `outcome='rejected_parse'` |

Load tests run in a staging env sized identically to production (RDS instance parity is mandatory). Chaos tests use `chaos-mesh` for pod kills and network partitions between EKS and RDS.

---

## 14. Eight-Week Implementation Plan

**Team: 6 engineers.** Roles: 2 backend (ingress/consumers), 1 DB (schema/migrations/perf), 1 SRE (obs/deploy/load), 1 security (IAM/KMS/replay), 1 tech lead (integration, reconciliation, cross-team).

### Workstreams and sequence

| Wk | Ingress | Consumers | DB & Data | SRE/Obs | Security | Gates |
|----|---------|-----------|-----------|---------|----------|-------|
| 1 | Extract merchant_account_id per provider; feature flag | — | Schema DDL + review; PgBouncer pilot | Metrics/log spec | KMS keys, IAM roles | Design review sign-off |
| 2 | Envelope v2 builder; unit tests | Consumer scaffold + effect_idem | Schema deployed to staging; partman | Dashboards v0 | S3 lifecycle policies | **G1**: schema in staging, ingress builds envelope |
| 3 | Shadow journaling in prod (write-only) | Shadow consumer skeleton | Load test #1 journal writes | Alerting v1 | Replay CLI stub | **G2**: journal error rate < 0.01% for 72h |
| 4 | Dual publish v1+v2 | Shadow decision logging | Reconcile shadow vs. live | Dashboards for shadow diff | Replay CLI approvals | **G3**: shadow decision parity ≥ 99.9% |
| 5 | — | Ledger v2 authoritative (ramp 1→100%) | Load test #2 full path | Ledger v2 dashboard | Audit log verification | **G4**: Ledger v2 stable 72h, 0 duplicate effects |
| 6 | — | Entitlements v2 authoritative | DDB recovery worker | Entitlements dashboard | — | **G5**: Entitlements v2 stable 72h |
| 7 | Publish-after-commit + janitor hardening | State machine polish; stalled sweeper | Reconciliation job daily | Reconciliation dashboard | Retention automation | **G6**: Reconciliation clean for 72h |
| 8 | Burn-in; chaos tests | Chaos + burst load | Partition archive tested | Runbooks | Pen test replay controls | **G7**: All success criteria met |

### Smallest safe first production scope (weeks 1–5)

If timeline slips, the minimum viable release is:
- Ingress with merchant_account_id extraction, journal, and canonical event.
- Ledger v2 authoritative.
- Entitlements **remains on v1** with a temporary duplicate-suppression patch (query `canonical_event` before applying); full v2 in a follow-up.
- Email and Analytics unchanged.

This eliminates the two highest-impact bugs (Ledger duplicate credits, merchant-collision suppression) in 5 weeks.

### Dependencies

- G2 blocks G3 (must have reliable journal before dual publish is meaningful).
- G3 blocks G4 (shadow parity is the safety net for Ledger cutover).
- G4 blocks G5 (Ledger cutover pattern is validated before repeating for Entitlements).
- Reconciliation (G6) blocks final sign-off.

---

## 15. Risk Register

| # | Risk | Likelihood | Impact | Mitigation | Measurable trigger |
|---|------|-----------|--------|------------|--------------------|
| R1 | Journal write latency exceeds budget under burst | Med | High | PgBouncer, batched inserts, load test #1; fallback to async journal with WAL-buffered guarantee only if load test proves synchronous impossible | Ingress p99 > 80 ms for 10 min |
| R2 | Idempotency key derivation bug creates collisions | Low | Very High | Per-provider unit tests, shadow validation, reconciliation alerts | `idempotency_key_collisions_suspected_total` > 0 |
| R3 | Synthesized keys drift when provider mutates fields | Med | Med | Field allowlist versioning; retry-count reconciliation | Synthetic-key duplicates > 0.1% |
| R4 | DynamoDB throttling under burst | Med | Med | On-demand capacity; recovery worker | `ThrottledRequests` > 100/min |
| R5 | Stuck `pending_events` buffer grows unbounded | Low | High | Hard cap 20, stalled sweeper, alerts | `payment_state_pending_events_gauge` > 5 |
| R6 | Migration rollback creates duplicate effects | Low | Very High | Effect_idem is enforced on both paths; ramped rollout | `effect_apply_conflict_total` rate change > 3σ during rollback |
| R7 | RDS writer saturation from workload growth | Med | High | Load tests, headroom monitoring, plan for read replicas / partitioning | CPU > 70% sustained or IOPS > 80% |
| R8 | Replay abused by insider | Low | High | Approval workflow, audit log, rate limit | Replay outside change window; > 1 replay per operator per day |
| R9 | Provider webhook contract changes | Med | Med | Contract tests per provider quarterly; parse failure alerts | `journal.outcome='rejected_parse'` > 10/min |
| R10 | Postgres unique-index bloat on canonical_event | Low | Med | Monitor index size; partition after 500M rows | Index size > 500 GB |
| R11 | Team velocity insufficient for 8 weeks | Med | Med | Defined minimum scope (§14); descope Entitlements to patch | Gate slip > 1 week |
| R12 | Payload minimization drops a field FinOps later needs | Med | Low | payload_digest retained; provider re-fetch possible; allowlist reviewed with FinOps | Reconciliation gap traceable to missing field |

---

## 16. Architecture Decision Record

**ADR-001: Payment integrity remediation architecture**

**Status**: Proposed

**Context**: See sections 0–2. Existing pipeline lacks idempotency, has a TOCTOU race in Ledger, uses the wrong collision domain for provider event IDs, and has no journal for FinOps audit.

**Decision**:
1. Introduce a Postgres-backed durable delivery journal written synchronously before HTTP 200.
2. Introduce a canonical event table with unique index on `(provider, merchant_account_id, provider_event_id_or_synthetic_hash)`.
3. Enforce business-effect idempotency via per-consumer `*_effect_idem` tables using `INSERT ... ON CONFLICT` co-committed with the effect in Postgres.
4. Bridge to DynamoDB using an outbox pattern in Postgres plus DynamoDB conditional writes keyed on the same `effect_id`.
5. Separate transport deduplication from business-state validation via an explicit payment lifecycle state machine with bounded buffering.
6. Keep SNS/SQS standard for fan-out; do not adopt FIFO.
7. Migrate with shadow journaling, dual publish, shadow consumer, and ramped cutover per consumer.

**Rejected alternatives**:
- **SNS FIFO + SQS FIFO with content dedup**: Insufficient dedup window (5 min vs. 7-day replay), does not solve collision-domain bug, does not remove need for consumer idempotency, and forces every consumer onto FIFO. §2.5.
- **Redis-based idempotency**: Constraint explicitly forbids Redis as source of truth. Also lacks the transactional co-commit with ledger writes.
- **Debezium CDC from Postgres to consumers**: Adds significant operational surface for marginal benefit over SNS publish-after-commit + janitor. Reserved as future option if publish latency becomes a bottleneck.
- **DynamoDB transactions with Postgres via 2PC**: No native support; XA over these systems is unreliable. Outbox pattern is the accepted trade-off.
- **Kinesis Data Streams for ingress**: Adds another primitive and re-sharding complexity; SNS fan-out to per-consumer SQS is simpler and matches existing operational knowledge.
- **Storing raw payloads in Postgres**: Violates security constraint on PII minimization.

**Consequences**:
- **Positive**: Duplicate business effects become structurally impossible (bounded by Postgres unique constraints and DynamoDB conditional writes). Merchant-collision bug eliminated. FinOps has a complete, immutable audit trail. Migration does not require coordinated deploys of every consumer.
- **Negative**: Ingress path now depends on Postgres availability. RDS becomes a critical dependency for webhook ACK. Mitigated by RDS Multi-AZ, PgBouncer, and load-tested capacity headroom; documented as a known risk (R1, R7).
- **Operational**: New reconciliation and replay workflows require FinOps training and runbook adoption.
- **Cost**: Additional RDS storage (~2 TB/mo hot, S3 archive thereafter), one new SNS topic during transition, no material DynamoDB cost change. Estimated < 8% increase in payments-platform infra cost.

**Review date**: 6 months post-launch, or on any P1 payment integrity incident.
Finish reason: stop
Try in Playground
import { createClient } from '@runware/sdk'

const client = await createClient({ apiKey: process.env.RUNWARE_API_KEY })
await client.connect()

const [result] = await client.run({
  model: 'anthropic:claude@opus-4.7',
  settings: {
    thinkingLevel: 'high'
  },
  messages: [
    {
      role: 'user',
      content: 'Act as the principal engineer responsible for resolving a recurring payment-integrity incident at Northstar Billing, a multi-tenant subscription platform. Produce a technical remediation plan that engineering, SRE, security, finance operations, and support can approve and execute.\n\nCURRENT SYSTEM\n- A Go API receives Stripe, Adyen, and PayPal webhooks through an AWS Application Load Balancer.\n- API instances validate provider signatures, translate payloads into an internal event shape, publish to an Amazon SNS standard topic, and immediately return HTTP 200.\n- SNS fans out to SQS standard queues for Ledger, Entitlements, Email, and Analytics services.\n- Ledger workers write to PostgreSQL 15. Entitlements workers update DynamoDB. Each service retries independently.\n- The internal event currently contains provider, provider event ID, tenant ID, event type, payment ID, amount, currency, provider timestamp, and payload.\n- There is no shared idempotency mechanism. Ledger attempts to detect duplicates by querying provider_event_id before inserting, but the check and insert are separate statements.\n- Some tenants use multiple merchant accounts. Provider event IDs are unique only within a merchant account, although merchant account ID is not currently included in the internal event.\n- Deployments run on EKS across three availability zones. PostgreSQL uses RDS Multi-AZ with one writer. Redis exists but is configured as a non-durable cache and must not be the source of truth.\n\nINCIDENT EVIDENCE\n- During a 47-minute SNS delivery degradation, retry bursts produced 18,420 duplicate deliveries from 6,130 distinct provider events.\n- 286 duplicate ledger entries were created. Of these, 41 triggered duplicate customer credits and 17 triggered duplicate entitlement extensions.\n- Two legitimate recurring-payment events were incorrectly suppressed because different merchant accounts emitted the same provider event ID.\n- Duplicate deliveries can arrive concurrently, minutes later, or up to seven days later after manual replay.\n- Events from a single payment can arrive out of order; for example, payment.refunded may be observed before payment.captured.\n- Finance requires an immutable record of every received delivery, including duplicates, signature result, original receipt time, and processing outcome.\n- Security prohibits storing full raw webhook payloads when they contain unnecessary personal data. Required financial records must remain available for seven years; operational delivery metadata may be archived after 90 days.\n\nBUSINESS AND DELIVERY CONSTRAINTS\n- Sustained traffic is 2,000 webhook deliveries per second, with bursts to 10,000 per second for five minutes.\n- Added webhook-ingress latency must remain below 80 ms at p99.\n- No more than five minutes of planned write unavailability is acceptable.\n- The migration must not pause webhook intake or require coordinated deployment of every consumer.\n- Existing consumers cannot all be changed immediately. Ledger and Entitlements can be modified in the first release; Email and Analytics will migrate later.\n- At-least-once transport is acceptable. Do not claim exactly-once delivery. The required outcome is effectively-once execution of each business effect.\n- PostgreSQL remains the ledger system of record. You may introduce additional AWS managed services only when justified.\n- The first production release must be achievable by six engineers within eight weeks.\n- Assume provider signatures are already implemented correctly; signature redesign is out of scope.\n\nREQUIRED DELIVERABLE\nWrite a self-contained engineering design document in Markdown. Use concrete decisions rather than presenting an unranked menu of options. Include:\n\n1. Executive summary and explicit success criteria.\n2. Failure analysis explaining the race condition, collision domain, retry behavior, out-of-order events, and why SNS/SQS FIFO alone would not solve the complete problem.\n3. The proposed target architecture from ingress through durable receipt, publication, consumption, business-effect execution, replay, and audit. Clearly identify every transactional boundary and source of truth.\n4. A canonical event envelope with field names, types, required/optional status, and semantics. Define a collision-safe idempotency key that accounts for provider and merchant account scope. Explain handling for providers that do not supply a stable event ID.\n5. PostgreSQL table designs for the immutable delivery journal, canonical events, processing attempts, and business-effect idempotency records. Provide representative PostgreSQL DDL, including primary keys, unique constraints, indexes, partitioning or retention strategy, and only the most important columns.\n6. Transaction pseudocode for both Ledger and Entitlements consumers. Show how concurrent duplicate deliveries are resolved atomically, how failures are retried safely, and how a claimed operation is recovered if a worker dies. Address the fact that DynamoDB cannot participate in the PostgreSQL transaction.\n7. A state-machine strategy for out-of-order payment lifecycle events. Distinguish transport deduplication from business-state transition validation, including how refunds observed before captures are handled without silently losing either event.\n8. Capacity reasoning using the supplied sustained and burst rates. Estimate daily delivery volume, propose partition sizing and hot-path indexes, identify likely bottlenecks, and state which assumptions require load testing. Approximate calculations are sufficient but must be shown.\n9. A zero-intake-downtime migration plan with phases, dual-write or shadow-validation behavior, backward compatibility for unmigrated consumers, backfill strategy, rollback criteria, and safeguards against creating effects twice during rollback.\n10. An observability plan with specific metrics, structured-log fields, traces, dashboards, and alert thresholds. Include detection of duplicates prevented, key collisions, invalid transitions, stuck work, replay activity, journal lag, and reconciliation mismatches.\n11. A reconciliation and repair procedure finance operations can run. Explain how it compares provider records, the delivery journal, canonical events, and ledger effects without automatically issuing money-moving corrections.\n12. A security, privacy, and retention section covering payload minimization, encryption, tenant isolation, privileged replay controls, seven-year financial retention, and archival or deletion of operational data.\n13. A focused test strategy covering concurrency, crash recovery, delayed duplicates, event-ID collisions across merchant accounts, out-of-order events, partial downstream failure, replay, migration compatibility, and burst load.\n14. An eight-week implementation sequence divided into workstreams, with dependencies, ownership by role, go/no-go gates, and the smallest safe first production scope.\n15. A risk register listing likelihood, impact, mitigation, and a measurable trigger for each major risk.\n16. A final architecture decision record summarizing the chosen approach, rejected alternatives, and consequences.\n\nBe precise about delivery semantics and do not use vague claims such as exactly-once processing without defining the scope and mechanism. Explicitly separate facts supplied above from assumptions you introduce. Keep the document detailed enough for implementation review while prioritizing the highest-risk decisions over generic background explanation.'
    }
  ]
})
import asyncio
import os

from runware import Runware


async def main():
    async with Runware(api_key=os.environ["RUNWARE_API_KEY"]) as client:
        results = await client.run({
            "model": "anthropic:claude@opus-4.7",
            "settings": {
                "thinkingLevel": "high"
            },
            "messages": [
                {
                    "role": "user",
                    "content": "Act as the principal engineer responsible for resolving a recurring payment-integrity incident at Northstar Billing, a multi-tenant subscription platform. Produce a technical remediation plan that engineering, SRE, security, finance operations, and support can approve and execute.\n\nCURRENT SYSTEM\n- A Go API receives Stripe, Adyen, and PayPal webhooks through an AWS Application Load Balancer.\n- API instances validate provider signatures, translate payloads into an internal event shape, publish to an Amazon SNS standard topic, and immediately return HTTP 200.\n- SNS fans out to SQS standard queues for Ledger, Entitlements, Email, and Analytics services.\n- Ledger workers write to PostgreSQL 15. Entitlements workers update DynamoDB. Each service retries independently.\n- The internal event currently contains provider, provider event ID, tenant ID, event type, payment ID, amount, currency, provider timestamp, and payload.\n- There is no shared idempotency mechanism. Ledger attempts to detect duplicates by querying provider_event_id before inserting, but the check and insert are separate statements.\n- Some tenants use multiple merchant accounts. Provider event IDs are unique only within a merchant account, although merchant account ID is not currently included in the internal event.\n- Deployments run on EKS across three availability zones. PostgreSQL uses RDS Multi-AZ with one writer. Redis exists but is configured as a non-durable cache and must not be the source of truth.\n\nINCIDENT EVIDENCE\n- During a 47-minute SNS delivery degradation, retry bursts produced 18,420 duplicate deliveries from 6,130 distinct provider events.\n- 286 duplicate ledger entries were created. Of these, 41 triggered duplicate customer credits and 17 triggered duplicate entitlement extensions.\n- Two legitimate recurring-payment events were incorrectly suppressed because different merchant accounts emitted the same provider event ID.\n- Duplicate deliveries can arrive concurrently, minutes later, or up to seven days later after manual replay.\n- Events from a single payment can arrive out of order; for example, payment.refunded may be observed before payment.captured.\n- Finance requires an immutable record of every received delivery, including duplicates, signature result, original receipt time, and processing outcome.\n- Security prohibits storing full raw webhook payloads when they contain unnecessary personal data. Required financial records must remain available for seven years; operational delivery metadata may be archived after 90 days.\n\nBUSINESS AND DELIVERY CONSTRAINTS\n- Sustained traffic is 2,000 webhook deliveries per second, with bursts to 10,000 per second for five minutes.\n- Added webhook-ingress latency must remain below 80 ms at p99.\n- No more than five minutes of planned write unavailability is acceptable.\n- The migration must not pause webhook intake or require coordinated deployment of every consumer.\n- Existing consumers cannot all be changed immediately. Ledger and Entitlements can be modified in the first release; Email and Analytics will migrate later.\n- At-least-once transport is acceptable. Do not claim exactly-once delivery. The required outcome is effectively-once execution of each business effect.\n- PostgreSQL remains the ledger system of record. You may introduce additional AWS managed services only when justified.\n- The first production release must be achievable by six engineers within eight weeks.\n- Assume provider signatures are already implemented correctly; signature redesign is out of scope.\n\nREQUIRED DELIVERABLE\nWrite a self-contained engineering design document in Markdown. Use concrete decisions rather than presenting an unranked menu of options. Include:\n\n1. Executive summary and explicit success criteria.\n2. Failure analysis explaining the race condition, collision domain, retry behavior, out-of-order events, and why SNS/SQS FIFO alone would not solve the complete problem.\n3. The proposed target architecture from ingress through durable receipt, publication, consumption, business-effect execution, replay, and audit. Clearly identify every transactional boundary and source of truth.\n4. A canonical event envelope with field names, types, required/optional status, and semantics. Define a collision-safe idempotency key that accounts for provider and merchant account scope. Explain handling for providers that do not supply a stable event ID.\n5. PostgreSQL table designs for the immutable delivery journal, canonical events, processing attempts, and business-effect idempotency records. Provide representative PostgreSQL DDL, including primary keys, unique constraints, indexes, partitioning or retention strategy, and only the most important columns.\n6. Transaction pseudocode for both Ledger and Entitlements consumers. Show how concurrent duplicate deliveries are resolved atomically, how failures are retried safely, and how a claimed operation is recovered if a worker dies. Address the fact that DynamoDB cannot participate in the PostgreSQL transaction.\n7. A state-machine strategy for out-of-order payment lifecycle events. Distinguish transport deduplication from business-state transition validation, including how refunds observed before captures are handled without silently losing either event.\n8. Capacity reasoning using the supplied sustained and burst rates. Estimate daily delivery volume, propose partition sizing and hot-path indexes, identify likely bottlenecks, and state which assumptions require load testing. Approximate calculations are sufficient but must be shown.\n9. A zero-intake-downtime migration plan with phases, dual-write or shadow-validation behavior, backward compatibility for unmigrated consumers, backfill strategy, rollback criteria, and safeguards against creating effects twice during rollback.\n10. An observability plan with specific metrics, structured-log fields, traces, dashboards, and alert thresholds. Include detection of duplicates prevented, key collisions, invalid transitions, stuck work, replay activity, journal lag, and reconciliation mismatches.\n11. A reconciliation and repair procedure finance operations can run. Explain how it compares provider records, the delivery journal, canonical events, and ledger effects without automatically issuing money-moving corrections.\n12. A security, privacy, and retention section covering payload minimization, encryption, tenant isolation, privileged replay controls, seven-year financial retention, and archival or deletion of operational data.\n13. A focused test strategy covering concurrency, crash recovery, delayed duplicates, event-ID collisions across merchant accounts, out-of-order events, partial downstream failure, replay, migration compatibility, and burst load.\n14. An eight-week implementation sequence divided into workstreams, with dependencies, ownership by role, go/no-go gates, and the smallest safe first production scope.\n15. A risk register listing likelihood, impact, mitigation, and a measurable trigger for each major risk.\n16. A final architecture decision record summarizing the chosen approach, rejected alternatives, and consequences.\n\nBe precise about delivery semantics and do not use vague claims such as exactly-once processing without defining the scope and mechanism. Explicitly separate facts supplied above from assumptions you introduce. Keep the document detailed enough for implementation review while prioritizing the highest-risk decisions over generic background explanation."
                }
            ]
        })


asyncio.run(main())
curl https://api.runware.ai/v1 \
  -H "Authorization: Bearer $RUNWARE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '[
    {
      "taskType": "textInference",
      "taskUUID": "b8a06f35-ec82-4f52-b555-67e02821f370",
      "model": "anthropic:claude@opus-4.7",
      "settings": {
        "thinkingLevel": "high"
      },
      "messages": [
        {
          "role": "user",
          "content": "Act as the principal engineer responsible for resolving a recurring payment-integrity incident at Northstar Billing, a multi-tenant subscription platform. Produce a technical remediation plan that engineering, SRE, security, finance operations, and support can approve and execute.\n\nCURRENT SYSTEM\n- A Go API receives Stripe, Adyen, and PayPal webhooks through an AWS Application Load Balancer.\n- API instances validate provider signatures, translate payloads into an internal event shape, publish to an Amazon SNS standard topic, and immediately return HTTP 200.\n- SNS fans out to SQS standard queues for Ledger, Entitlements, Email, and Analytics services.\n- Ledger workers write to PostgreSQL 15. Entitlements workers update DynamoDB. Each service retries independently.\n- The internal event currently contains provider, provider event ID, tenant ID, event type, payment ID, amount, currency, provider timestamp, and payload.\n- There is no shared idempotency mechanism. Ledger attempts to detect duplicates by querying provider_event_id before inserting, but the check and insert are separate statements.\n- Some tenants use multiple merchant accounts. Provider event IDs are unique only within a merchant account, although merchant account ID is not currently included in the internal event.\n- Deployments run on EKS across three availability zones. PostgreSQL uses RDS Multi-AZ with one writer. Redis exists but is configured as a non-durable cache and must not be the source of truth.\n\nINCIDENT EVIDENCE\n- During a 47-minute SNS delivery degradation, retry bursts produced 18,420 duplicate deliveries from 6,130 distinct provider events.\n- 286 duplicate ledger entries were created. Of these, 41 triggered duplicate customer credits and 17 triggered duplicate entitlement extensions.\n- Two legitimate recurring-payment events were incorrectly suppressed because different merchant accounts emitted the same provider event ID.\n- Duplicate deliveries can arrive concurrently, minutes later, or up to seven days later after manual replay.\n- Events from a single payment can arrive out of order; for example, payment.refunded may be observed before payment.captured.\n- Finance requires an immutable record of every received delivery, including duplicates, signature result, original receipt time, and processing outcome.\n- Security prohibits storing full raw webhook payloads when they contain unnecessary personal data. Required financial records must remain available for seven years; operational delivery metadata may be archived after 90 days.\n\nBUSINESS AND DELIVERY CONSTRAINTS\n- Sustained traffic is 2,000 webhook deliveries per second, with bursts to 10,000 per second for five minutes.\n- Added webhook-ingress latency must remain below 80 ms at p99.\n- No more than five minutes of planned write unavailability is acceptable.\n- The migration must not pause webhook intake or require coordinated deployment of every consumer.\n- Existing consumers cannot all be changed immediately. Ledger and Entitlements can be modified in the first release; Email and Analytics will migrate later.\n- At-least-once transport is acceptable. Do not claim exactly-once delivery. The required outcome is effectively-once execution of each business effect.\n- PostgreSQL remains the ledger system of record. You may introduce additional AWS managed services only when justified.\n- The first production release must be achievable by six engineers within eight weeks.\n- Assume provider signatures are already implemented correctly; signature redesign is out of scope.\n\nREQUIRED DELIVERABLE\nWrite a self-contained engineering design document in Markdown. Use concrete decisions rather than presenting an unranked menu of options. Include:\n\n1. Executive summary and explicit success criteria.\n2. Failure analysis explaining the race condition, collision domain, retry behavior, out-of-order events, and why SNS/SQS FIFO alone would not solve the complete problem.\n3. The proposed target architecture from ingress through durable receipt, publication, consumption, business-effect execution, replay, and audit. Clearly identify every transactional boundary and source of truth.\n4. A canonical event envelope with field names, types, required/optional status, and semantics. Define a collision-safe idempotency key that accounts for provider and merchant account scope. Explain handling for providers that do not supply a stable event ID.\n5. PostgreSQL table designs for the immutable delivery journal, canonical events, processing attempts, and business-effect idempotency records. Provide representative PostgreSQL DDL, including primary keys, unique constraints, indexes, partitioning or retention strategy, and only the most important columns.\n6. Transaction pseudocode for both Ledger and Entitlements consumers. Show how concurrent duplicate deliveries are resolved atomically, how failures are retried safely, and how a claimed operation is recovered if a worker dies. Address the fact that DynamoDB cannot participate in the PostgreSQL transaction.\n7. A state-machine strategy for out-of-order payment lifecycle events. Distinguish transport deduplication from business-state transition validation, including how refunds observed before captures are handled without silently losing either event.\n8. Capacity reasoning using the supplied sustained and burst rates. Estimate daily delivery volume, propose partition sizing and hot-path indexes, identify likely bottlenecks, and state which assumptions require load testing. Approximate calculations are sufficient but must be shown.\n9. A zero-intake-downtime migration plan with phases, dual-write or shadow-validation behavior, backward compatibility for unmigrated consumers, backfill strategy, rollback criteria, and safeguards against creating effects twice during rollback.\n10. An observability plan with specific metrics, structured-log fields, traces, dashboards, and alert thresholds. Include detection of duplicates prevented, key collisions, invalid transitions, stuck work, replay activity, journal lag, and reconciliation mismatches.\n11. A reconciliation and repair procedure finance operations can run. Explain how it compares provider records, the delivery journal, canonical events, and ledger effects without automatically issuing money-moving corrections.\n12. A security, privacy, and retention section covering payload minimization, encryption, tenant isolation, privileged replay controls, seven-year financial retention, and archival or deletion of operational data.\n13. A focused test strategy covering concurrency, crash recovery, delayed duplicates, event-ID collisions across merchant accounts, out-of-order events, partial downstream failure, replay, migration compatibility, and burst load.\n14. An eight-week implementation sequence divided into workstreams, with dependencies, ownership by role, go/no-go gates, and the smallest safe first production scope.\n15. A risk register listing likelihood, impact, mitigation, and a measurable trigger for each major risk.\n16. A final architecture decision record summarizing the chosen approach, rejected alternatives, and consequences.\n\nBe precise about delivery semantics and do not use vague claims such as exactly-once processing without defining the scope and mechanism. Explicitly separate facts supplied above from assumptions you introduce. Keep the document detailed enough for implementation review while prioritizing the highest-risk decisions over generic background explanation."
        }
      ]
    }
  ]'
runware run anthropic:claude@opus-4.7 \
  settings.thinkingLevel=high \
  messages.0.role=user \
  messages.0.content="Act as the principal engineer responsible for resolving a recurring payment-integrity incident at Northstar Billing, a multi-tenant subscription platform. Produce a technical remediation plan that engineering, SRE, security, finance operations, and support can approve and execute.

CURRENT SYSTEM
- A Go API receives Stripe, Adyen, and PayPal webhooks through an AWS Application Load Balancer.
- API instances validate provider signatures, translate payloads into an internal event shape, publish to an Amazon SNS standard topic, and immediately return HTTP 200.
- SNS fans out to SQS standard queues for Ledger, Entitlements, Email, and Analytics services.
- Ledger workers write to PostgreSQL 15. Entitlements workers update DynamoDB. Each service retries independently.
- The internal event currently contains provider, provider event ID, tenant ID, event type, payment ID, amount, currency, provider timestamp, and payload.
- There is no shared idempotency mechanism. Ledger attempts to detect duplicates by querying provider_event_id before inserting, but the check and insert are separate statements.
- Some tenants use multiple merchant accounts. Provider event IDs are unique only within a merchant account, although merchant account ID is not currently included in the internal event.
- Deployments run on EKS across three availability zones. PostgreSQL uses RDS Multi-AZ with one writer. Redis exists but is configured as a non-durable cache and must not be the source of truth.

INCIDENT EVIDENCE
- During a 47-minute SNS delivery degradation, retry bursts produced 18,420 duplicate deliveries from 6,130 distinct provider events.
- 286 duplicate ledger entries were created. Of these, 41 triggered duplicate customer credits and 17 triggered duplicate entitlement extensions.
- Two legitimate recurring-payment events were incorrectly suppressed because different merchant accounts emitted the same provider event ID.
- Duplicate deliveries can arrive concurrently, minutes later, or up to seven days later after manual replay.
- Events from a single payment can arrive out of order; for example, payment.refunded may be observed before payment.captured.
- Finance requires an immutable record of every received delivery, including duplicates, signature result, original receipt time, and processing outcome.
- Security prohibits storing full raw webhook payloads when they contain unnecessary personal data. Required financial records must remain available for seven years; operational delivery metadata may be archived after 90 days.

BUSINESS AND DELIVERY CONSTRAINTS
- Sustained traffic is 2,000 webhook deliveries per second, with bursts to 10,000 per second for five minutes.
- Added webhook-ingress latency must remain below 80 ms at p99.
- No more than five minutes of planned write unavailability is acceptable.
- The migration must not pause webhook intake or require coordinated deployment of every consumer.
- Existing consumers cannot all be changed immediately. Ledger and Entitlements can be modified in the first release; Email and Analytics will migrate later.
- At-least-once transport is acceptable. Do not claim exactly-once delivery. The required outcome is effectively-once execution of each business effect.
- PostgreSQL remains the ledger system of record. You may introduce additional AWS managed services only when justified.
- The first production release must be achievable by six engineers within eight weeks.
- Assume provider signatures are already implemented correctly; signature redesign is out of scope.

REQUIRED DELIVERABLE
Write a self-contained engineering design document in Markdown. Use concrete decisions rather than presenting an unranked menu of options. Include:

1. Executive summary and explicit success criteria.
2. Failure analysis explaining the race condition, collision domain, retry behavior, out-of-order events, and why SNS/SQS FIFO alone would not solve the complete problem.
3. The proposed target architecture from ingress through durable receipt, publication, consumption, business-effect execution, replay, and audit. Clearly identify every transactional boundary and source of truth.
4. A canonical event envelope with field names, types, required/optional status, and semantics. Define a collision-safe idempotency key that accounts for provider and merchant account scope. Explain handling for providers that do not supply a stable event ID.
5. PostgreSQL table designs for the immutable delivery journal, canonical events, processing attempts, and business-effect idempotency records. Provide representative PostgreSQL DDL, including primary keys, unique constraints, indexes, partitioning or retention strategy, and only the most important columns.
6. Transaction pseudocode for both Ledger and Entitlements consumers. Show how concurrent duplicate deliveries are resolved atomically, how failures are retried safely, and how a claimed operation is recovered if a worker dies. Address the fact that DynamoDB cannot participate in the PostgreSQL transaction.
7. A state-machine strategy for out-of-order payment lifecycle events. Distinguish transport deduplication from business-state transition validation, including how refunds observed before captures are handled without silently losing either event.
8. Capacity reasoning using the supplied sustained and burst rates. Estimate daily delivery volume, propose partition sizing and hot-path indexes, identify likely bottlenecks, and state which assumptions require load testing. Approximate calculations are sufficient but must be shown.
9. A zero-intake-downtime migration plan with phases, dual-write or shadow-validation behavior, backward compatibility for unmigrated consumers, backfill strategy, rollback criteria, and safeguards against creating effects twice during rollback.
10. An observability plan with specific metrics, structured-log fields, traces, dashboards, and alert thresholds. Include detection of duplicates prevented, key collisions, invalid transitions, stuck work, replay activity, journal lag, and reconciliation mismatches.
11. A reconciliation and repair procedure finance operations can run. Explain how it compares provider records, the delivery journal, canonical events, and ledger effects without automatically issuing money-moving corrections.
12. A security, privacy, and retention section covering payload minimization, encryption, tenant isolation, privileged replay controls, seven-year financial retention, and archival or deletion of operational data.
13. A focused test strategy covering concurrency, crash recovery, delayed duplicates, event-ID collisions across merchant accounts, out-of-order events, partial downstream failure, replay, migration compatibility, and burst load.
14. An eight-week implementation sequence divided into workstreams, with dependencies, ownership by role, go/no-go gates, and the smallest safe first production scope.
15. A risk register listing likelihood, impact, mitigation, and a measurable trigger for each major risk.
16. A final architecture decision record summarizing the chosen approach, rejected alternatives, and consequences.

Be precise about delivery semantics and do not use vague claims such as exactly-once processing without defining the scope and mechanism. Explicitly separate facts supplied above from assumptions you introduce. Keep the document detailed enough for implementation review while prioritizing the highest-risk decisions over generic background explanation."
{
  "taskType": "textInference",
  "taskUUID": "b8a06f35-ec82-4f52-b555-67e02821f370",
  "model": "anthropic:claude@opus-4.7",
  "settings": {
    "thinkingLevel": "high"
  },
  "messages": [
    {
      "role": "user",
      "content": "Act as the principal engineer responsible for resolving a recurring payment-integrity incident at Northstar Billing, a multi-tenant subscription platform. Produce a technical remediation plan that engineering, SRE, security, finance operations, and support can approve and execute.\n\nCURRENT SYSTEM\n- A Go API receives Stripe, Adyen, and PayPal webhooks through an AWS Application Load Balancer.\n- API instances validate provider signatures, translate payloads into an internal event shape, publish to an Amazon SNS standard topic, and immediately return HTTP 200.\n- SNS fans out to SQS standard queues for Ledger, Entitlements, Email, and Analytics services.\n- Ledger workers write to PostgreSQL 15. Entitlements workers update DynamoDB. Each service retries independently.\n- The internal event currently contains provider, provider event ID, tenant ID, event type, payment ID, amount, currency, provider timestamp, and payload.\n- There is no shared idempotency mechanism. Ledger attempts to detect duplicates by querying provider_event_id before inserting, but the check and insert are separate statements.\n- Some tenants use multiple merchant accounts. Provider event IDs are unique only within a merchant account, although merchant account ID is not currently included in the internal event.\n- Deployments run on EKS across three availability zones. PostgreSQL uses RDS Multi-AZ with one writer. Redis exists but is configured as a non-durable cache and must not be the source of truth.\n\nINCIDENT EVIDENCE\n- During a 47-minute SNS delivery degradation, retry bursts produced 18,420 duplicate deliveries from 6,130 distinct provider events.\n- 286 duplicate ledger entries were created. Of these, 41 triggered duplicate customer credits and 17 triggered duplicate entitlement extensions.\n- Two legitimate recurring-payment events were incorrectly suppressed because different merchant accounts emitted the same provider event ID.\n- Duplicate deliveries can arrive concurrently, minutes later, or up to seven days later after manual replay.\n- Events from a single payment can arrive out of order; for example, payment.refunded may be observed before payment.captured.\n- Finance requires an immutable record of every received delivery, including duplicates, signature result, original receipt time, and processing outcome.\n- Security prohibits storing full raw webhook payloads when they contain unnecessary personal data. Required financial records must remain available for seven years; operational delivery metadata may be archived after 90 days.\n\nBUSINESS AND DELIVERY CONSTRAINTS\n- Sustained traffic is 2,000 webhook deliveries per second, with bursts to 10,000 per second for five minutes.\n- Added webhook-ingress latency must remain below 80 ms at p99.\n- No more than five minutes of planned write unavailability is acceptable.\n- The migration must not pause webhook intake or require coordinated deployment of every consumer.\n- Existing consumers cannot all be changed immediately. Ledger and Entitlements can be modified in the first release; Email and Analytics will migrate later.\n- At-least-once transport is acceptable. Do not claim exactly-once delivery. The required outcome is effectively-once execution of each business effect.\n- PostgreSQL remains the ledger system of record. You may introduce additional AWS managed services only when justified.\n- The first production release must be achievable by six engineers within eight weeks.\n- Assume provider signatures are already implemented correctly; signature redesign is out of scope.\n\nREQUIRED DELIVERABLE\nWrite a self-contained engineering design document in Markdown. Use concrete decisions rather than presenting an unranked menu of options. Include:\n\n1. Executive summary and explicit success criteria.\n2. Failure analysis explaining the race condition, collision domain, retry behavior, out-of-order events, and why SNS/SQS FIFO alone would not solve the complete problem.\n3. The proposed target architecture from ingress through durable receipt, publication, consumption, business-effect execution, replay, and audit. Clearly identify every transactional boundary and source of truth.\n4. A canonical event envelope with field names, types, required/optional status, and semantics. Define a collision-safe idempotency key that accounts for provider and merchant account scope. Explain handling for providers that do not supply a stable event ID.\n5. PostgreSQL table designs for the immutable delivery journal, canonical events, processing attempts, and business-effect idempotency records. Provide representative PostgreSQL DDL, including primary keys, unique constraints, indexes, partitioning or retention strategy, and only the most important columns.\n6. Transaction pseudocode for both Ledger and Entitlements consumers. Show how concurrent duplicate deliveries are resolved atomically, how failures are retried safely, and how a claimed operation is recovered if a worker dies. Address the fact that DynamoDB cannot participate in the PostgreSQL transaction.\n7. A state-machine strategy for out-of-order payment lifecycle events. Distinguish transport deduplication from business-state transition validation, including how refunds observed before captures are handled without silently losing either event.\n8. Capacity reasoning using the supplied sustained and burst rates. Estimate daily delivery volume, propose partition sizing and hot-path indexes, identify likely bottlenecks, and state which assumptions require load testing. Approximate calculations are sufficient but must be shown.\n9. A zero-intake-downtime migration plan with phases, dual-write or shadow-validation behavior, backward compatibility for unmigrated consumers, backfill strategy, rollback criteria, and safeguards against creating effects twice during rollback.\n10. An observability plan with specific metrics, structured-log fields, traces, dashboards, and alert thresholds. Include detection of duplicates prevented, key collisions, invalid transitions, stuck work, replay activity, journal lag, and reconciliation mismatches.\n11. A reconciliation and repair procedure finance operations can run. Explain how it compares provider records, the delivery journal, canonical events, and ledger effects without automatically issuing money-moving corrections.\n12. A security, privacy, and retention section covering payload minimization, encryption, tenant isolation, privileged replay controls, seven-year financial retention, and archival or deletion of operational data.\n13. A focused test strategy covering concurrency, crash recovery, delayed duplicates, event-ID collisions across merchant accounts, out-of-order events, partial downstream failure, replay, migration compatibility, and burst load.\n14. An eight-week implementation sequence divided into workstreams, with dependencies, ownership by role, go/no-go gates, and the smallest safe first production scope.\n15. A risk register listing likelihood, impact, mitigation, and a measurable trigger for each major risk.\n16. A final architecture decision record summarizing the chosen approach, rejected alternatives, and consequences.\n\nBe precise about delivery semantics and do not use vague claims such as exactly-once processing without defining the scope and mechanism. Explicitly separate facts supplied above from assumptions you introduce. Keep the document detailed enough for implementation review while prioritizing the highest-risk decisions over generic background explanation."
    }
  ]
}
Response
{
  "taskType": "textInference",
  "taskUUID": "b8a06f35-ec82-4f52-b555-67e02821f370",
  "text": "# Northstar Billing — Payment Integrity Remediation Design\n\n**Status:** Proposed, pending sign-off (Engineering, SRE, Security, FinOps, Support)\n**Author:** Principal Engineer, Payments Platform\n**Version:** 1.0\n\n---\n\n## 0. Terminology and Provenance\n\n**Delivery semantics used throughout:**\n- **Transport delivery**: at-least-once. Explicitly not deduplicated by SNS/SQS.\n- **Ingress persistence**: one durable journal row per received delivery, including duplicates.\n- **Canonical event materialization**: exactly one canonical row per idempotency key (enforced by unique index in PostgreSQL).\n- **Business-effect execution**: **effectively-once** — each business effect (ledger posting, entitlement mutation, email send) is applied at most once per canonical event via a per-effect idempotency record co-committed with the effect (PostgreSQL) or via a conditional write keyed on the same identifier (DynamoDB).\n\nWe do **not** claim exactly-once end-to-end. We claim: (a) at-least-once transport, (b) exactly-once canonicalization, (c) effectively-once business effects.\n\n**Assumptions introduced (not supplied in brief), flagged where used:**\n- **A1**: Adyen `pspReference` + `eventCode` is stable per merchant account; PayPal `id` (event resource) is stable per merchant; Stripe `evt_...` is stable per Stripe account. Requires verification during Phase 0.\n- **A2**: Merchant account identifier can be derived from webhook path or payload (`account` header for Stripe, `merchantAccountCode` for Adyen, `resource.supplementary_data.related_ids` for PayPal). Ingress must be extended to extract it. Verified per provider.\n- **A3**: Payload compressed size (gzip) averages 2 KB, max 32 KB. Requires measurement.\n- **A4**: RDS instance is db.r6g.4xlarge or larger with provisioned IOPS. If not, sizing in §8 changes.\n- **A5**: Existing signature validation cannot be bypassed to accept a delivery; a signature failure produces a journal row with `signature_valid=false` and no canonical event.\n\n---\n\n## 1. Executive Summary and Success Criteria\n\n### 1.1 Summary\n\nThe current pipeline treats webhook receipt, deduplication, and business-effect execution as loosely coupled operations with no shared idempotency contract. Ledger's \"SELECT then INSERT\" duplicate check is a classic time-of-check/time-of-use race; DynamoDB updates in Entitlements have no idempotency guard at all; and the internal event lacks merchant account scope, causing legitimate events to collide.\n\nThe remediation introduces:\n\n1. **A durable delivery journal** in PostgreSQL, written synchronously before ACK-ing the provider (source of truth for \"we received it\").\n2. **A canonical event table** with a **collision-safe idempotency key** unique on `(provider, merchant_account_id, provider_event_id_or_hash)`, materialized once per event.\n3. **Per-consumer business-effect idempotency records** co-committed with the effect in PostgreSQL, and **conditional writes** keyed by the same effect ID in DynamoDB.\n4. **A payment-lifecycle state machine** in the Ledger that distinguishes transport duplication (dropped) from out-of-order state transitions (buffered and reconciled, never silently discarded).\n5. **A migration path** that dual-writes and shadow-validates without pausing intake.\n\n### 1.2 Success Criteria (measurable)\n\n| # | Criterion | Target | Measurement |\n|---|-----------|--------|-------------|\n| S1 | Duplicate business effects per 10⁶ deliveries | 0 (bounded by DB uniqueness) | Monthly reconciliation report |\n| S2 | Legitimate events dropped due to key collision | 0 | Reconciliation vs. provider dashboards |\n| S3 | Added ingress latency p99 | ≤ 80 ms | Ingress histogram |\n| S4 | Planned write downtime during migration | ≤ 5 min | Change record |\n| S5 | Journal completeness (deliveries recorded / deliveries acknowledged) | 100.000% | Journal count vs. ALB access logs |\n| S6 | Time to replay a single event, controlled | ≤ 60 s from operator command | Runbook drill |\n| S7 | Reconciliation report generated per tenant per day | ≥ 99% success | Batch job SLO |\n\n---\n\n## 2. Failure Analysis\n\n### 2.1 The race condition in Ledger\n\n```\nT1: SELECT 1 FROM ledger WHERE provider_event_id = 'evt_X'  -> 0 rows\nT2: SELECT 1 FROM ledger WHERE provider_event_id = 'evt_X'  -> 0 rows\nT1: INSERT ... 'evt_X'  -> OK\nT2: INSERT ... 'evt_X'  -> OK   ← duplicate\n```\nNo unique constraint, no `SELECT ... FOR UPDATE`, no `INSERT ... ON CONFLICT`. Under SNS retry bursts, concurrent consumers observe the same pre-insert state. This produced the 286 duplicate ledger entries.\n\n### 2.2 Collision domain error\n\nProvider event IDs are unique **within a merchant account**, not globally. Ledger's dedup key was `provider_event_id` alone. When two merchant accounts emitted `evt_1a2b3c` for legitimate distinct events, the second was suppressed as a \"duplicate.\" Root cause: `merchant_account_id` was never in the internal event.\n\n### 2.3 Retry behavior\n\nSNS standard + SQS standard: at-least-once with unbounded duplicate window in practice (SQS visibility timeouts, DLQ replays, manual replay up to 7 days). During the 47-minute degradation, SNS retried aggressively; each retry was a new SQS message. There is no message-ID-level dedup across retries.\n\n### 2.4 Out-of-order events\n\nPayment lifecycle (`authorized → captured → refunded`) can be observed in any order because: (a) providers emit events on independent workers, (b) SNS/SQS standard reorder freely, (c) retries of an early event may arrive after a later event succeeded. A naive \"apply in arrival order\" ledger accepts a refund against a payment that has no capture record, then rejects the capture as invalid — silently losing state.\n\n### 2.5 Why SNS/SQS FIFO alone is insufficient\n\nFIFO with content-based dedup would fix **transport duplicates within a 5-minute window** for a single `MessageGroupId`. It does not solve:\n\n1. **7-day manual replay** exceeds FIFO's 5-minute dedup window.\n2. **Consumer-side idempotency**: FIFO delivery is still at-least-once (visibility timeout expiry re-delivers). Business-effect dedup still required.\n3. **Ordering scope**: FIFO orders within a `MessageGroupId`. To preserve payment ordering, group must be `payment_id`, but that serializes all events per payment and destroys the fan-out throughput (10k/s burst → thousands of groups is fine, but a hot payment blocks). Also SNS→SQS FIFO does not fan out to standard queues, forcing every consumer to FIFO.\n4. **Cross-consumer idempotency**: FIFO dedup is per-queue. Ledger and Entitlements would each need independent guarantees.\n5. **Collision domain**: FIFO dedup uses message content hash. It does not know that provider event IDs are only unique per merchant.\n6. **Journal requirement**: FinOps needs *every* delivery recorded, including duplicates. FIFO discarding duplicates before the journal would violate this.\n\n**Decision**: keep SNS/SQS standard for fan-out; solve idempotency and ordering in PostgreSQL where the constraints are enforceable and auditable.\n\n---\n\n## 3. Target Architecture\n\n### 3.1 Data flow\n\n```\nProvider → ALB → Ingress API (Go)\n                    │\n                    ├─(1) verify signature\n                    ├─(2) extract merchant_account_id, build envelope\n                    ├─(3) BEGIN tx\n                    │      INSERT delivery_journal (always)\n                    │      INSERT canonical_event  ON CONFLICT DO NOTHING\n                    │    COMMIT\n                    ├─(4) if canonical row was new → publish to SNS\n                    └─(5) return HTTP 200\n\nSNS topic (standard, fan-out)\n  ├→ SQS: ledger.q         → Ledger workers → PostgreSQL (ledger + effect_idem in same tx)\n  ├→ SQS: entitlements.q   → Entitlements workers → DynamoDB (conditional write) + effect_idem in Postgres\n  ├→ SQS: email.q          → (unchanged in phase 1, tolerates duplicates)\n  └→ SQS: analytics.q      → (unchanged in phase 1)\n\nAsync:\n  - Journal archiver → S3 (Parquet, partitioned by date) after 90 days\n  - Reconciliation job → S3 report per tenant per day\n  - Replay CLI → publishes canonical events back onto SNS with replay flag\n```\n\n### 3.2 Transactional boundaries and sources of truth\n\n| Concern | Source of truth | Transactional mechanism |\n|---|---|---|\n| \"Did we receive this delivery?\" | `delivery_journal` (Postgres) | Single-row insert, committed before HTTP 200 |\n| \"Have we canonicalized this event?\" | `canonical_event` (Postgres) | Unique index on idempotency key |\n| \"Has this ledger effect been applied?\" | `ledger_effect_idem` + `ledger_entry` (Postgres) | Same transaction |\n| \"Has this entitlement effect been applied?\" | `entitlement_effect_idem` (Postgres) is authoritative; DynamoDB is projection | Postgres commit first, then DynamoDB conditional write; recovery loop reconciles |\n| \"Current payment lifecycle state\" | `payment_state` (Postgres) | Row-level lock + state-transition validation |\n| Cache | Redis | Never source of truth; may be flushed |\n\n### 3.3 Why the journal-before-ACK trade-off is acceptable\n\nAdding a single Postgres insert to the ingress path costs ~3–8 ms p50, ~20–40 ms p99 on a well-sized RDS with connection pooling (PgBouncer transaction mode). Well within the 80 ms budget. The alternative — journal asynchronously — risks losing the audit record if the API crashes between ACK and journal, which FinOps has ruled out.\n\n### 3.4 Publish-after-commit\n\nSNS publish happens **after** the Postgres commit. If SNS publish fails, the canonical row exists but no message is emitted. A janitor process scans `canonical_event WHERE published_at IS NULL AND created_at < now() - interval '30 seconds'` and re-publishes. This is safe because consumers are idempotent on canonical event ID.\n\n---\n\n## 4. Canonical Event Envelope\n\n### 4.1 Envelope (JSON, versioned)\n\n| Field | Type | Req | Semantics |\n|---|---|---|---|\n| `schema_version` | string (semver) | Y | Envelope schema, e.g. `\"1.0\"` |\n| `event_id` | UUIDv7 | Y | Northstar-assigned; primary key of `canonical_event` |\n| `idempotency_key` | string (≤ 200 chars) | Y | See §4.2 |\n| `provider` | enum(`stripe`,`adyen`,`paypal`) | Y | |\n| `merchant_account_id` | string | Y | Provider-scoped account (Stripe account, Adyen merchant, PayPal payer) |\n| `tenant_id` | UUID | Y | Northstar tenant |\n| `provider_event_id` | string | N | Provider's ID if supplied; null if provider omits |\n| `event_type` | string | Y | Canonicalized, e.g. `payment.captured` |\n| `payment_id` | string | N | Provider payment/charge ID; required for lifecycle events |\n| `amount_minor` | int64 | N | Minor units (cents); required for money-moving events |\n| `currency` | ISO-4217 string | N | Required with `amount_minor` |\n| `provider_event_time` | RFC3339 | Y | From provider |\n| `received_at` | RFC3339 | Y | Ingress receipt time (server clock) |\n| `payload_ref` | string | Y | S3 URI + SHA-256 of minimized payload; see §12 |\n| `payload_digest` | hex(SHA-256) | Y | Hash of full raw payload (for audit, without storing it) |\n| `signature_valid` | bool | Y | Result of provider signature check |\n| `producer_version` | string | Y | Ingress build hash |\n\n`payload` is **not** in the envelope on the wire to downstream consumers; they fetch minimized fields via `payload_ref` if needed. Reduces SNS message size, PII exposure, and cost.\n\n### 4.2 Idempotency key\n\n```\nidempotency_key = provider + \":\" + merchant_account_id + \":\" + event_scope\n\nwhere event_scope =\n  if provider_event_id present:  provider_event_id\n  else:                          \"syn:\" + sha256_hex(\n                                    payment_id || event_type ||\n                                    provider_event_time_iso ||\n                                    amount_minor || currency ||\n                                    canonical_provider_specific_fields\n                                 )\n```\n\n- **Collision safety**: including `merchant_account_id` fixes the observed suppression bug.\n- **Missing provider event ID (synthesized keys)**: PayPal `TRANSACTION.SALE.COMPLETED` under some flows lacks a stable event ID; synthesize from stable business fields. `syn:` prefix makes provenance explicit and reconciliable.\n- **Length**: ≤ 200 chars fits Stripe (`evt_` + 24), Adyen (`pspReference` up to 16 + eventCode), PayPal IDs, plus prefix.\n- **Stability**: derived only from provider-supplied stable fields; never includes our timestamps or UUIDs.\n\n### 4.3 Handling providers with no stable ID\n\nSynthesized keys are inherently weaker: if the provider re-sends the same event with the same fields, we correctly dedupe; if they mutate a field on retry (rare, but Adyen has been observed to), we would double-count. Mitigation:\n\n1. Per-provider allowlist of fields used in synthesis, reviewed quarterly.\n2. Reconciliation report flags synthetic-key events with retry counts > 1 for manual review.\n3. Prefer provider event ID everywhere it exists.\n\n---\n\n## 5. PostgreSQL Schema\n\n### 5.1 Delivery journal (append-only, immutable)\n\n```sql\nCREATE TABLE delivery_journal (\n    delivery_id         UUID        NOT NULL DEFAULT gen_random_uuid(),\n    received_at         TIMESTAMPTZ NOT NULL,\n    provider            TEXT        NOT NULL,\n    merchant_account_id TEXT,\n    provider_event_id   TEXT,\n    tenant_id           UUID,\n    signature_valid     BOOLEAN     NOT NULL,\n    payload_digest      BYTEA       NOT NULL,          -- sha256 of raw body\n    payload_bytes       INT         NOT NULL,\n    source_ip           INET,\n    http_headers_digest BYTEA,\n    idempotency_key     TEXT,                          -- null if signature invalid or unparseable\n    canonical_event_id  UUID,                          -- null if not canonicalized\n    outcome             TEXT NOT NULL,                 -- 'accepted_new','accepted_duplicate','rejected_signature','rejected_parse'\n    PRIMARY KEY (received_at, delivery_id)\n) PARTITION BY RANGE (received_at);\n\n-- Monthly partitions; pg_partman manages creation and detach.\nCREATE INDEX ON delivery_journal (idempotency_key) WHERE idempotency_key IS NOT NULL;\nCREATE INDEX ON delivery_journal (canonical_event_id) WHERE canonical_event_id IS NOT NULL;\nCREATE INDEX ON delivery_journal (provider, received_at);\n```\n\n**Retention**: Operational metadata may be archived after 90 days (per brief). Journal partitions older than 90 days are exported to S3 (Parquet, KMS-encrypted, Object Lock in compliance mode for 7 years) then detached and dropped. The S3 export **is** the 7-year record for delivery metadata. Money-moving records live in `ledger_entry` (7-year retention in Postgres + logical backups).\n\n### 5.2 Canonical events\n\n```sql\nCREATE TABLE canonical_event (\n    event_id            UUID        PRIMARY KEY,       -- UUIDv7\n    idempotency_key     TEXT        NOT NULL,\n    provider            TEXT        NOT NULL,\n    merchant_account_id TEXT        NOT NULL,\n    tenant_id           UUID        NOT NULL,\n    event_type          TEXT        NOT NULL,\n    payment_id          TEXT,\n    amount_minor        BIGINT,\n    currency            CHAR(3),\n    provider_event_time TIMESTAMPTZ NOT NULL,\n    received_at         TIMESTAMPTZ NOT NULL,\n    payload_ref         TEXT        NOT NULL,\n    schema_version      TEXT        NOT NULL,\n    published_at        TIMESTAMPTZ,\n    created_at          TIMESTAMPTZ NOT NULL DEFAULT now()\n);\n\nCREATE UNIQUE INDEX canonical_event_idem\n    ON canonical_event (idempotency_key);\nCREATE INDEX canonical_event_payment\n    ON canonical_event (tenant_id, payment_id)\n    WHERE payment_id IS NOT NULL;\nCREATE INDEX canonical_event_publish_pending\n    ON canonical_event (created_at)\n    WHERE published_at IS NULL;\n```\n\nNot partitioned initially; sized in §8. Add range partitioning by `created_at` when row count exceeds ~500M (est. ~3 years).\n\n### 5.3 Processing attempts\n\n```sql\nCREATE TABLE processing_attempt (\n    attempt_id      BIGSERIAL PRIMARY KEY,\n    event_id        UUID        NOT NULL REFERENCES canonical_event(event_id),\n    consumer        TEXT        NOT NULL,            -- 'ledger','entitlements',...\n    worker_id       TEXT        NOT NULL,\n    started_at      TIMESTAMPTZ NOT NULL DEFAULT now(),\n    finished_at     TIMESTAMPTZ,\n    outcome         TEXT,                            -- 'applied','skipped_duplicate','deferred','failed'\n    error_class     TEXT,\n    error_message   TEXT\n);\nCREATE INDEX ON processing_attempt (event_id, consumer);\nCREATE INDEX ON processing_attempt (consumer, started_at)\n    WHERE finished_at IS NULL;                       -- stuck-work detection\n```\n\n### 5.4 Per-consumer business-effect idempotency\n\n```sql\nCREATE TABLE ledger_effect_idem (\n    effect_id       TEXT        PRIMARY KEY,          -- see §6.1\n    event_id        UUID        NOT NULL REFERENCES canonical_event(event_id),\n    ledger_entry_id UUID        NOT NULL,             -- FK to ledger_entry\n    applied_at      TIMESTAMPTZ NOT NULL DEFAULT now()\n);\n\nCREATE TABLE entitlement_effect_idem (\n    effect_id       TEXT        PRIMARY KEY,\n    event_id        UUID        NOT NULL REFERENCES canonical_event(event_id),\n    dynamo_pk       TEXT        NOT NULL,\n    dynamo_sk       TEXT        NOT NULL,\n    intent          JSONB       NOT NULL,             -- what we intended to apply\n    ddb_confirmed_at TIMESTAMPTZ,                     -- set after DynamoDB confirms\n    applied_at      TIMESTAMPTZ NOT NULL DEFAULT now()\n);\nCREATE INDEX ON entitlement_effect_idem (ddb_confirmed_at)\n    WHERE ddb_confirmed_at IS NULL;                   -- recovery worker scans\n```\n\n### 5.5 Payment lifecycle state\n\n```sql\nCREATE TABLE payment_state (\n    tenant_id           UUID  NOT NULL,\n    provider            TEXT  NOT NULL,\n    merchant_account_id TEXT  NOT NULL,\n    payment_id          TEXT  NOT NULL,\n    state               TEXT  NOT NULL,               -- 'unknown','authorized','captured','refunded','voided','failed'\n    last_event_time     TIMESTAMPTZ NOT NULL,\n    pending_events      JSONB NOT NULL DEFAULT '[]',  -- buffered out-of-order events\n    updated_at          TIMESTAMPTZ NOT NULL DEFAULT now(),\n    PRIMARY KEY (tenant_id, provider, merchant_account_id, payment_id)\n);\n```\n\n---\n\n## 6. Consumer Transaction Logic\n\n### 6.1 Effect ID construction\n\n```\neffect_id = consumer_name + \":\" + event_id + \":\" + effect_kind\n```\nE.g. `ledger:018f...:posting`, `entitlements:018f...:extend`. Deterministic from canonical event; safe as primary key.\n\n### 6.2 Ledger consumer pseudocode\n\n```\non_message(msg):\n    env = parse(msg)\n    tx = pg.begin(SERIALIZABLE)          # or READ COMMITTED with ON CONFLICT; see below\n    try:\n        # 1. Claim the work with a processing_attempt row\n        attempt_id = tx.insert(processing_attempt, {\n            event_id: env.event_id, consumer:'ledger',\n            worker_id: WORKER_ID, started_at: now()\n        })\n\n        # 2. Advisory lock scoped to payment to serialize state transitions\n        tx.exec(\"SELECT pg_advisory_xact_lock(hashtextextended($1,0))\",\n                env.tenant_id || env.payment_id)\n\n        # 3. Load and validate lifecycle transition\n        state = tx.select_for_update(payment_state, keys...)\n        decision = lifecycle.evaluate(state, env)     # apply | buffer | ignore_duplicate\n        if decision == ignore_duplicate:\n            tx.update(processing_attempt, id=attempt_id,\n                      outcome='skipped_duplicate', finished_at=now())\n            tx.commit(); ack(msg); return\n\n        if decision == buffer:\n            state.pending_events = append(state.pending_events, env.minified())\n            tx.update(payment_state, state)\n            tx.update(processing_attempt, id=attempt_id,\n                      outcome='deferred', finished_at=now())\n            tx.commit(); ack(msg); return            # will be re-drained when predecessor arrives\n\n        # decision == apply: also drain any pending events now applicable\n        events_to_apply = [env] + lifecycle.drain(state, env)\n\n        for e in events_to_apply:\n            effect_id = \"ledger:\" + e.event_id + \":posting\"\n            # 4. Atomic effect: idem-row + ledger row in one tx\n            r = tx.exec(\"\"\"\n                INSERT INTO ledger_effect_idem(effect_id, event_id, ledger_entry_id)\n                VALUES ($1, $2, $3)\n                ON CONFLICT (effect_id) DO NOTHING\n                RETURNING effect_id\n            \"\"\", effect_id, e.event_id, new_uuid)\n            if r.rows == 0:\n                continue                              # already applied\n            tx.insert(ledger_entry, ...)              # references ledger_entry_id above\n            state = lifecycle.advance(state, e)\n\n        tx.update(payment_state, state)\n        tx.update(processing_attempt, id=attempt_id,\n                  outcome='applied', finished_at=now())\n        tx.commit()\n        ack(msg)\n    except SerializationFailure:\n        tx.rollback(); nack(msg)                      # SQS visibility timeout will redeliver\n    except Exception as e:\n        tx.rollback()\n        record_failed_attempt(env.event_id, e)\n        nack(msg)                                     # DLQ after N attempts\n```\n\n**Concurrency correctness**: `pg_advisory_xact_lock` on `payment_id` serializes per-payment work. `INSERT ... ON CONFLICT DO NOTHING` on `ledger_effect_idem` (primary key) is atomic even without the advisory lock; the lock exists to serialize state-machine transitions, not idempotency.\n\n**Worker death recovery**: An orphaned `processing_attempt` with `finished_at IS NULL` older than 10 minutes is swept by a reaper that (a) checks whether the effect row exists — if yes, marks attempt `applied`; if no, marks `failed` and re-enqueues by publishing to the retry SNS topic. Since the effect row is the source of truth, the reaper cannot cause duplicates.\n\n### 6.3 Entitlements consumer pseudocode\n\nDynamoDB cannot participate in the Postgres transaction. Pattern: **outbox in Postgres, then apply, then confirm.**\n\n```\non_message(msg):\n    env = parse(msg)\n    effect_id = \"entitlements:\" + env.event_id + \":extend\"\n    intent = build_intent(env)                        # deterministic\n\n    # Phase A: reserve the intent atomically in Postgres\n    tx = pg.begin()\n    r = tx.exec(\"\"\"\n        INSERT INTO entitlement_effect_idem(effect_id, event_id, dynamo_pk, dynamo_sk, intent)\n        VALUES ($1, $2, $3, $4, $5)\n        ON CONFLICT (effect_id) DO NOTHING\n        RETURNING effect_id, ddb_confirmed_at\n    \"\"\", effect_id, env.event_id, intent.pk, intent.sk, intent)\n    if r.rows == 0:\n        row = tx.select_one(entitlement_effect_idem, effect_id=effect_id)\n        tx.commit()\n        if row.ddb_confirmed_at is not None:\n            ack(msg); return                          # fully done\n        # else: fall through to Phase B (retrying an unconfirmed prior attempt)\n    else:\n        tx.commit()\n\n    # Phase B: apply to DynamoDB with a conditional write keyed on effect_id\n    try:\n        ddb.update_item(\n            TableName='entitlements',\n            Key={'pk': intent.pk, 'sk': intent.sk},\n            UpdateExpression='ADD applied_effects :eid SET expires_at = :exp',\n            ConditionExpression='NOT contains(applied_effects, :eid)',\n            ExpressionAttributeValues={\n                ':eid': effect_id,\n                ':exp': intent.new_expiry,\n            },\n        )\n    except ConditionalCheckFailedException:\n        pass                                          # already applied to DDB previously\n\n    # Phase C: confirm\n    pg.exec(\"UPDATE entitlement_effect_idem SET ddb_confirmed_at = now() WHERE effect_id = $1\",\n            effect_id)\n    ack(msg)\n```\n\n**Correctness**:\n- If crash after Phase A, before Phase B: next redelivery finds `ddb_confirmed_at IS NULL`, retries Phase B. DynamoDB conditional write is idempotent on `effect_id`.\n- If crash after Phase B, before Phase C: next redelivery retries Phase B; conditional check fails harmlessly; Phase C runs.\n- If DynamoDB accepts write but Postgres update fails: effect exists in DDB, next redelivery sees unconfirmed idem row, re-attempts DDB (idempotent), confirms in Postgres.\n- Recovery worker scans `entitlement_effect_idem WHERE ddb_confirmed_at IS NULL AND applied_at < now() - interval '5 minutes'` and drives Phase B+C without an inbound message.\n\nThe `applied_effects` set in DynamoDB has bounded growth per entitlement item; a TTL/GC job trims effect IDs older than 30 days once the corresponding Postgres row is confirmed.\n\n---\n\n## 7. Out-of-Order Lifecycle State Machine\n\n### 7.1 Two concerns, separated\n\n| Concern | Mechanism | Table |\n|---|---|---|\n| Transport duplicate | Same idempotency key already canonicalized | `canonical_event` unique index |\n| Business duplicate effect | Same `effect_id` already applied | `*_effect_idem` |\n| Out-of-order transition | State machine defers, buffers, and later drains | `payment_state.pending_events` |\n\nTransport dedup **must not** silently discard out-of-order events. A refund arriving before its capture is not a duplicate — it is a valid future event.\n\n### 7.2 State machine\n\nStates: `unknown → authorized → captured → refunded (terminal) | voided (terminal) | failed (terminal)`\n\nTransition table (excerpt):\n\n| From \\ Event | authorized | captured | refunded | voided | failed |\n|---|---|---|---|---|---|\n| unknown | → authorized | **buffer** | **buffer** | **buffer** | → failed |\n| authorized | ignore (dup) | → captured | **buffer** | → voided | → failed |\n| captured | ignore | ignore | → refunded | reject-invalid | reject-invalid |\n| refunded | ignore | ignore | ignore/partial-refund handling | reject | reject |\n\n`buffer` = write into `payment_state.pending_events`; do not apply ledger effect yet; ACK message (canonical event is durable, effect will materialize when the predecessor arrives).\n\n`ignore` = we've seen this event class already; treat as duplicate (but still record `processing_attempt` with `outcome='skipped_duplicate'`).\n\n`reject-invalid` = state contradiction. Do **not** silently drop. Route to `payment_state_conflict` table for FinOps review, alert fires.\n\n### 7.3 Buffering bound\n\n`pending_events` capped at 20 entries per payment (way beyond expected 3–5). If exceeded, alert fires: indicates provider misbehavior or missing predecessor event that requires manual intervention. Bounded buffer prevents unbounded row growth.\n\n### 7.4 Time-based drain\n\nA background sweeper runs every 5 minutes: for `payment_state` rows with `pending_events` non-empty and `updated_at < now() - interval '1 hour'`, it emits a `payment_state_stalled` metric per (tenant, provider) and creates a Support ticket for the specific payment. Prevents silent losses when the predecessor event never arrives.\n\n---\n\n## 8. Capacity Reasoning\n\n### 8.1 Volumes\n\n- Sustained: 2,000 req/s → **172.8M deliveries/day**.\n- Burst: 10,000 req/s × 5 min = 3M deliveries in 5 min.\n- Assumed duplicate multiplier under normal ops: 1.05×. Under incident: up to 3× (matches evidence: 18,420 / 6,130 ≈ 3.0).\n\n### 8.2 Journal write load\n\n- Sustained: 2,000 inserts/s into `delivery_journal`. Each row ~250 bytes + 2 indexes ≈ 400 B written per row after WAL amplification.\n- Burst: 10,000 inserts/s → ~4 MB/s write, ~40 MB/s WAL with fsync amplification. Well within db.r6g.4xlarge with gp3 12,000 IOPS.\n- Monthly partition size at sustained: 172.8M × 30 × 250 B ≈ **1.3 TB/month** raw; ~2 TB with indexes. Partition + archive to S3 after 90 days keeps hot storage at ~6 TB.\n\n**Requires load test**: sustained 10k/s insert into partitioned parent with 2 indexes on db.r6g.4xlarge. If p99 insert latency exceeds 30 ms, upgrade to r6g.8xlarge or shard journal by `provider` prefix (separate tables per provider, UNION view for reads).\n\n### 8.3 Canonical event and effect writes\n\n- Sustained ~1,900/s new canonical events (5% duplicates suppressed).\n- Ledger effect insert rate: ~1,900/s. Effect table growth: 172.8M × 0.95 × ~150 B ≈ 25 GB/month. Manageable without partitioning for ~2 years.\n\n### 8.4 Ingress latency budget\n\nPath (target p99):\n\n| Stage | Budget (ms) |\n|---|---|\n| TLS + ALB | 10 |\n| Signature verify | 8 |\n| Parse + envelope build | 5 |\n| PgBouncer + tx (insert × 2, commit) | 35 |\n| SNS publish (async, not on critical path if commit-then-publish outside request) | — |\n| Slack | 22 |\n| **Total** | **80** |\n\n**Decision**: Return HTTP 200 after Postgres commit but **before** SNS publish. SNS publish happens on a background goroutine with retry, and the janitor process backstops failures. This keeps SNS latency (~15 ms p99, spike-prone) out of the provider-facing path.\n\n### 8.5 Bottlenecks and mitigations\n\n| Bottleneck | Symptom | Mitigation |\n|---|---|---|\n| RDS writer IOPS under burst | Journal insert p99 spike | Provisioned IOPS, PgBouncer transaction pool, batch inserts (100-row COPY per 20ms window in ingress) |\n| Advisory-lock contention on hot payment | Ledger throughput drops for that payment | Acceptable; state transitions must serialize per payment |\n| DynamoDB throttling on entitlement update | Phase B retries pile up | On-demand capacity; monitor `ThrottledRequests` |\n| Postgres connection storm from EKS pods | `too many connections` | PgBouncer at ~1000 client → 200 server connections |\n\n**Requires load test**: burst behavior, PgBouncer sizing, batched journal insert efficacy.\n\n---\n\n## 9. Migration Plan (Zero-Intake-Downtime)\n\n### 9.1 Principles\n\n- Never pause webhook intake.\n- Dual-write during transition; canonical event is authoritative once shadow validation passes.\n- Every phase has an explicit rollback that cannot produce duplicate effects.\n\n### 9.2 Phases\n\n**Phase 0 — Foundations (week 1–2)**\n- Deploy new schema (journal, canonical_event, effect_idem, payment_state) to RDS with `CREATE TABLE IF NOT EXISTS`. No writes yet.\n- Deploy PgBouncer.\n- Extend ingress to extract `merchant_account_id` per provider; ship behind a feature flag with logging only.\n\n**Phase 1 — Shadow journaling (week 3)**\n- Ingress writes `delivery_journal` and `canonical_event` in shadow mode: rows written, SNS still uses the old envelope.\n- Metric: `journal_write_error_rate`. Gate: < 0.01% for 72 hours.\n\n**Phase 2 — Dual publish (week 4)**\n- Ingress publishes both old envelope (unchanged topic) and new envelope (new SNS topic `payment-events-v2`).\n- Ledger and Entitlements subscribe to v2 in **shadow mode**: they compute what they *would* do and log to `shadow_decision` table, without applying effects.\n- Gate: reconcile shadow decisions vs. actual old-path decisions for 72 hours. Discrepancies triaged manually.\n\n**Phase 3 — Ledger cutover (week 5)**\n- Ledger v2 becomes authoritative; old ledger consumer stops. Feature-flagged per tenant, ramped 1% → 10% → 100% over 4 days.\n- Effect idempotency prevents duplicate application even if a message is processed by both old and new consumers during ramp — this is the critical safety property.\n- Rollback: disable v2 flag; old consumer resumes. Because effect_idem is checked by both (old consumer patched to read effect_idem on rollback), no duplicates.\n\n**Phase 4 — Entitlements cutover (week 6)**\n- Same pattern. DynamoDB conditional-write keyed on `effect_id` is the safety net.\n\n**Phase 5 — Email & Analytics** (post-8-weeks; out of scope for first release)\n- They continue consuming the v1 topic via a bridge that republishes canonical events in the old shape. When migrated, bridge is removed.\n\n### 9.3 Backfill\n\nBackfill of pre-migration data is **not** performed. Historical events remain in the legacy ledger. Canonical event table starts fresh at Phase 1. Reconciliation from Phase 1 forward.\n\n### 9.4 Rollback criteria\n\nAutomatic rollback triggers:\n- Ingress p99 latency > 100 ms for 10 minutes.\n- `journal_write_error_rate` > 0.1% for 5 minutes.\n- Canonical event uniqueness violations > 0.5% of inserts (indicates key derivation bug).\n- Ledger v2 producing effects at > 1.5× v1 rate on same input stream (shadow discrepancy).\n\nRollback playbook: flip flag → verify old consumer resumes → confirm no gap in journal → post-mortem.\n\n---\n\n## 10. Observability\n\n### 10.1 Metrics (Prometheus names)\n\n| Metric | Type | Labels | Alert |\n|---|---|---|---|\n| `ingress_deliveries_total` | counter | provider, outcome | — |\n| `ingress_latency_seconds` | histogram | provider | p99 > 80ms 5m |\n| `journal_insert_errors_total` | counter | error_class | > 10/min |\n| `canonical_event_duplicates_prevented_total` | counter | provider | Report only |\n| `canonical_event_publish_pending` | gauge | — | > 1000 for 5m |\n| `idempotency_key_collisions_suspected_total` | counter | provider | > 0 pages FinOps |\n| `consumer_processed_total` | counter | consumer, outcome | — |\n| `consumer_processing_seconds` | histogram | consumer | p99 > 2s 10m |\n| `effect_apply_conflict_total` | counter | consumer | Rate change > 3σ |\n| `payment_state_invalid_transition_total` | counter | provider, from_state, event | > 0 pages |\n| `payment_state_pending_events_gauge` | gauge | tenant | > 5 per payment |\n| `payment_state_stalled_total` | counter | provider | > 0/hour pages |\n| `entitlement_ddb_unconfirmed_gauge` | gauge | — | > 100 for 10m |\n| `processing_attempt_stuck_gauge` | gauge | consumer | > 50 for 10m |\n| `replay_events_total` | counter | operator, reason | Report + audit log |\n| `reconciliation_mismatch_total` | counter | tenant, category | > 0 opens ticket |\n\n### 10.2 Structured log fields (mandatory on every log line in this pipeline)\n\n`trace_id, span_id, delivery_id, event_id, idempotency_key, provider, merchant_account_id, tenant_id, payment_id (if any), consumer, outcome`\n\n### 10.3 Traces\n\nOpenTelemetry spans: `ingress.receive` → `ingress.verify` → `ingress.persist` → `ingress.publish` (async) → `consumer.claim` → `consumer.state_eval` → `consumer.apply` → `consumer.ack`. Propagate trace ID in envelope metadata.\n\n### 10.4 Dashboards\n\n- **Ingress health**: RPS, p50/p99 latency, journal error rate, signature failure rate.\n- **Deduplication**: duplicates prevented per hour, ratio vs. total, per provider.\n- **Consumer health**: throughput, effect apply rate, DLQ depth, stuck attempts.\n- **Lifecycle**: state distribution, invalid transitions, pending events by age.\n- **Reconciliation**: last successful run per tenant, open mismatches, aged mismatches.\n\n### 10.5 Alerts (priority)\n\nP1 pages: journal write errors > 0.1%, ingress p99 > 100ms sustained, `payment_state_invalid_transition_total` non-zero, `reconciliation_mismatch_total` for money-moving categories.\nP2 tickets: DLQ depth > 100, unconfirmed DDB effects > 100, replay activity outside change window.\n\n---\n\n## 11. Reconciliation and Repair\n\n### 11.1 Daily reconciliation job\n\nRuns per tenant per day. Reads:\n1. Provider records via API (Stripe `events.list`, Adyen report, PayPal reports) for date D.\n2. `delivery_journal` for date D, filtered to that tenant.\n3. `canonical_event` for date D.\n4. `ledger_entry` for date D.\n\nProduces a report with four sections:\n\n| Section | Definition | Action |\n|---|---|---|\n| **A. Provider without journal** | Event exists in provider but not in journal | Investigate — likely missed webhook. May require manual replay. |\n| **B. Journal without canonical** | Delivery accepted but not canonicalized | Investigate parse/signature bug. |\n| **C. Canonical without ledger effect** (money-moving types) | Consumer failed to apply | Check DLQ, `processing_attempt`. |\n| **D. Ledger effect without provider record** | Ghost entries | High severity; potential double-counting from pre-migration data. |\n\n### 11.2 No automatic money movement\n\nThe job **never** posts corrections. It generates a signed report to S3, an entry in `reconciliation_report`, and if any section is non-empty, opens a Jira ticket assigned to FinOps. Corrections require:\n\n1. FinOps analyst reviews report.\n2. Two-person approval in the FinOps repair tool.\n3. Repair emits an event through the same canonical pipeline with `event_type='ops.manual_correction'` and a distinct `effect_id` prefix. This appears in reconciliation on future runs, closing the loop.\n\n### 11.3 Comparison keys\n\nReconciliation joins on `(provider, merchant_account_id, provider_event_id)` where available; falls back to `(payment_id, event_type, amount_minor, currency, provider_event_time)` for synthetic-key events. Discrepancies in the synthetic space are treated as suspected collisions and reviewed.\n\n---\n\n## 12. Security, Privacy, Retention\n\n### 12.1 Payload minimization\n\nRaw payloads are **not** stored in Postgres. On ingress:\n1. Compute `payload_digest = sha256(raw_body)`; store in journal.\n2. Extract only allowlisted fields per provider into a minimized JSON (`payment_id`, `amount`, `currency`, `event_type`, `merchant_account_id`, non-PII metadata).\n3. Store minimized JSON in S3 (`s3://northstar-events/{tenant}/{yyyy/mm/dd}/{event_id}.json.gz`), SSE-KMS with per-tenant CMK.\n4. `payload_ref` in canonical event is the S3 URI.\n\nFull raw payload is retained transiently in the ingress process for signature verification, then discarded. No PII (card holder name, email, address) is persisted in Postgres.\n\n### 12.2 Encryption\n\n- RDS: encryption at rest (KMS), TLS in transit, client cert auth for consumer roles.\n- S3: SSE-KMS with per-tenant CMK for minimized payloads; SSE-KMS with platform CMK for journal archive.\n- DynamoDB: encryption at rest, VPC endpoints only.\n- SNS/SQS: SSE-KMS.\n\n### 12.3 Tenant isolation\n\n- Row-level security on all tables keyed by `tenant_id`; consumer roles carry tenant context via `SET LOCAL app.tenant_id`.\n- S3 prefix per tenant, IAM condition on prefix.\n- DynamoDB partition key includes `tenant_id`.\n\n### 12.4 Replay controls\n\nReplay CLI requires:\n- SSO with hardware key.\n- Second reviewer approval (`replay-approver` group) for any replay > 1 event.\n- Audit log to CloudTrail + `replay_audit` table with operator, reason, event IDs, ticket link.\n- Hard limit: 10,000 events per replay command; larger replays require change ticket.\n- Replays cannot bypass idempotency; effect_idem prevents double application.\n\n### 12.5 Retention\n\n| Data | Retention | Location | Mechanism |\n|---|---|---|---|\n| `ledger_entry` and related financial records | 7 years | Postgres + S3 backup | Retention policy documented, deletion requires FinOps + Legal |\n| Delivery journal (hot) | 90 days | Postgres partitions | Auto-detach and archive |\n| Delivery journal (archive) | 7 years | S3 Object Lock (compliance mode) | Immutable |\n| Canonical event | 7 years | Postgres (partitioned after year 1) | |\n| Minimized payloads | 7 years | S3 with lifecycle to Glacier at 90d | |\n| Processing attempts | 90 days | Postgres | Truncated older partitions |\n| Reconciliation reports | 7 years | S3 Object Lock | |\n\nDeletions of PII per GDPR/CCPA operate on the minimized payload S3 objects and non-financial fields; financial records are retained under the legal-basis exemption and documented in the DPA.\n\n---\n\n## 13. Test Strategy\n\n| Category | Test | Pass criteria |\n|---|---|---|\n| **Concurrency** | Fire 50 goroutines inserting same idempotency_key | Exactly 1 canonical row; 49 conflict errors handled |\n| **Concurrency** | Fire 50 parallel Ledger consumers on same event | Exactly 1 `ledger_effect_idem` row, 1 `ledger_entry` |\n| **Crash recovery** | Kill Ledger pod between insert and commit | Redelivery reprocesses; no duplicate; attempt reaper cleans stale row |\n| **Crash recovery** | Kill Entitlements between Phase A and Phase B | DDB gets exactly one update, `ddb_confirmed_at` set |\n| **Crash recovery** | Kill Entitlements between Phase B and Phase C | Idem row remains unconfirmed; recovery worker completes it; no duplicate DDB write |\n| **Delayed duplicates** | Replay a 6-day-old event | Duplicate detected via idempotency_key; journal has 2 delivery rows, 1 canonical row |\n| **Collision** | Two merchant accounts send `provider_event_id='evt_1'` for different payments | Both processed; distinct canonical rows |\n| **Out-of-order** | Send refund before capture | Refund buffered; capture applied; then buffered refund drained; ledger has both entries in correct order |\n| **Out-of-order** | Send refund with no capture ever arriving | After 1 hour, `payment_state_stalled` fires; ticket opened; refund not silently lost |\n| **Partial failure** | DynamoDB throws for 10 min | Retries with backoff; recovery worker drives to completion; no duplicate |\n| **Replay** | Operator replays 1,000 events already applied | 0 new effects; `replay_audit` populated; metrics show duplicates prevented |\n| **Migration compat** | Old and new consumers both active on same event | Exactly 1 effect (proven by effect_idem constraint) |\n| **Burst load** | 10k/s for 5 min, 10% synthetic duplicates | Ingress p99 < 80 ms; 0 message loss; journal complete |\n| **Signature failure** | Send invalid signature | Journal row with `signature_valid=false`, no canonical event, no downstream propagation |\n| **Fuzz** | Malformed payloads, oversized bodies, truncated JSON | Ingress rejects with 4xx, journal row `outcome='rejected_parse'` |\n\nLoad tests run in a staging env sized identically to production (RDS instance parity is mandatory). Chaos tests use `chaos-mesh` for pod kills and network partitions between EKS and RDS.\n\n---\n\n## 14. Eight-Week Implementation Plan\n\n**Team: 6 engineers.** Roles: 2 backend (ingress/consumers), 1 DB (schema/migrations/perf), 1 SRE (obs/deploy/load), 1 security (IAM/KMS/replay), 1 tech lead (integration, reconciliation, cross-team).\n\n### Workstreams and sequence\n\n| Wk | Ingress | Consumers | DB & Data | SRE/Obs | Security | Gates |\n|----|---------|-----------|-----------|---------|----------|-------|\n| 1 | Extract merchant_account_id per provider; feature flag | — | Schema DDL + review; PgBouncer pilot | Metrics/log spec | KMS keys, IAM roles | Design review sign-off |\n| 2 | Envelope v2 builder; unit tests | Consumer scaffold + effect_idem | Schema deployed to staging; partman | Dashboards v0 | S3 lifecycle policies | **G1**: schema in staging, ingress builds envelope |\n| 3 | Shadow journaling in prod (write-only) | Shadow consumer skeleton | Load test #1 journal writes | Alerting v1 | Replay CLI stub | **G2**: journal error rate < 0.01% for 72h |\n| 4 | Dual publish v1+v2 | Shadow decision logging | Reconcile shadow vs. live | Dashboards for shadow diff | Replay CLI approvals | **G3**: shadow decision parity ≥ 99.9% |\n| 5 | — | Ledger v2 authoritative (ramp 1→100%) | Load test #2 full path | Ledger v2 dashboard | Audit log verification | **G4**: Ledger v2 stable 72h, 0 duplicate effects |\n| 6 | — | Entitlements v2 authoritative | DDB recovery worker | Entitlements dashboard | — | **G5**: Entitlements v2 stable 72h |\n| 7 | Publish-after-commit + janitor hardening | State machine polish; stalled sweeper | Reconciliation job daily | Reconciliation dashboard | Retention automation | **G6**: Reconciliation clean for 72h |\n| 8 | Burn-in; chaos tests | Chaos + burst load | Partition archive tested | Runbooks | Pen test replay controls | **G7**: All success criteria met |\n\n### Smallest safe first production scope (weeks 1–5)\n\nIf timeline slips, the minimum viable release is:\n- Ingress with merchant_account_id extraction, journal, and canonical event.\n- Ledger v2 authoritative.\n- Entitlements **remains on v1** with a temporary duplicate-suppression patch (query `canonical_event` before applying); full v2 in a follow-up.\n- Email and Analytics unchanged.\n\nThis eliminates the two highest-impact bugs (Ledger duplicate credits, merchant-collision suppression) in 5 weeks.\n\n### Dependencies\n\n- G2 blocks G3 (must have reliable journal before dual publish is meaningful).\n- G3 blocks G4 (shadow parity is the safety net for Ledger cutover).\n- G4 blocks G5 (Ledger cutover pattern is validated before repeating for Entitlements).\n- Reconciliation (G6) blocks final sign-off.\n\n---\n\n## 15. Risk Register\n\n| # | Risk | Likelihood | Impact | Mitigation | Measurable trigger |\n|---|------|-----------|--------|------------|--------------------|\n| R1 | Journal write latency exceeds budget under burst | Med | High | PgBouncer, batched inserts, load test #1; fallback to async journal with WAL-buffered guarantee only if load test proves synchronous impossible | Ingress p99 > 80 ms for 10 min |\n| R2 | Idempotency key derivation bug creates collisions | Low | Very High | Per-provider unit tests, shadow validation, reconciliation alerts | `idempotency_key_collisions_suspected_total` > 0 |\n| R3 | Synthesized keys drift when provider mutates fields | Med | Med | Field allowlist versioning; retry-count reconciliation | Synthetic-key duplicates > 0.1% |\n| R4 | DynamoDB throttling under burst | Med | Med | On-demand capacity; recovery worker | `ThrottledRequests` > 100/min |\n| R5 | Stuck `pending_events` buffer grows unbounded | Low | High | Hard cap 20, stalled sweeper, alerts | `payment_state_pending_events_gauge` > 5 |\n| R6 | Migration rollback creates duplicate effects | Low | Very High | Effect_idem is enforced on both paths; ramped rollout | `effect_apply_conflict_total` rate change > 3σ during rollback |\n| R7 | RDS writer saturation from workload growth | Med | High | Load tests, headroom monitoring, plan for read replicas / partitioning | CPU > 70% sustained or IOPS > 80% |\n| R8 | Replay abused by insider | Low | High | Approval workflow, audit log, rate limit | Replay outside change window; > 1 replay per operator per day |\n| R9 | Provider webhook contract changes | Med | Med | Contract tests per provider quarterly; parse failure alerts | `journal.outcome='rejected_parse'` > 10/min |\n| R10 | Postgres unique-index bloat on canonical_event | Low | Med | Monitor index size; partition after 500M rows | Index size > 500 GB |\n| R11 | Team velocity insufficient for 8 weeks | Med | Med | Defined minimum scope (§14); descope Entitlements to patch | Gate slip > 1 week |\n| R12 | Payload minimization drops a field FinOps later needs | Med | Low | payload_digest retained; provider re-fetch possible; allowlist reviewed with FinOps | Reconciliation gap traceable to missing field |\n\n---\n\n## 16. Architecture Decision Record\n\n**ADR-001: Payment integrity remediation architecture**\n\n**Status**: Proposed\n\n**Context**: See sections 0–2. Existing pipeline lacks idempotency, has a TOCTOU race in Ledger, uses the wrong collision domain for provider event IDs, and has no journal for FinOps audit.\n\n**Decision**:\n1. Introduce a Postgres-backed durable delivery journal written synchronously before HTTP 200.\n2. Introduce a canonical event table with unique index on `(provider, merchant_account_id, provider_event_id_or_synthetic_hash)`.\n3. Enforce business-effect idempotency via per-consumer `*_effect_idem` tables using `INSERT ... ON CONFLICT` co-committed with the effect in Postgres.\n4. Bridge to DynamoDB using an outbox pattern in Postgres plus DynamoDB conditional writes keyed on the same `effect_id`.\n5. Separate transport deduplication from business-state validation via an explicit payment lifecycle state machine with bounded buffering.\n6. Keep SNS/SQS standard for fan-out; do not adopt FIFO.\n7. Migrate with shadow journaling, dual publish, shadow consumer, and ramped cutover per consumer.\n\n**Rejected alternatives**:\n- **SNS FIFO + SQS FIFO with content dedup**: Insufficient dedup window (5 min vs. 7-day replay), does not solve collision-domain bug, does not remove need for consumer idempotency, and forces every consumer onto FIFO. §2.5.\n- **Redis-based idempotency**: Constraint explicitly forbids Redis as source of truth. Also lacks the transactional co-commit with ledger writes.\n- **Debezium CDC from Postgres to consumers**: Adds significant operational surface for marginal benefit over SNS publish-after-commit + janitor. Reserved as future option if publish latency becomes a bottleneck.\n- **DynamoDB transactions with Postgres via 2PC**: No native support; XA over these systems is unreliable. Outbox pattern is the accepted trade-off.\n- **Kinesis Data Streams for ingress**: Adds another primitive and re-sharding complexity; SNS fan-out to per-consumer SQS is simpler and matches existing operational knowledge.\n- **Storing raw payloads in Postgres**: Violates security constraint on PII minimization.\n\n**Consequences**:\n- **Positive**: Duplicate business effects become structurally impossible (bounded by Postgres unique constraints and DynamoDB conditional writes). Merchant-collision bug eliminated. FinOps has a complete, immutable audit trail. Migration does not require coordinated deploys of every consumer.\n- **Negative**: Ingress path now depends on Postgres availability. RDS becomes a critical dependency for webhook ACK. Mitigated by RDS Multi-AZ, PgBouncer, and load-tested capacity headroom; documented as a known risk (R1, R7).\n- **Operational**: New reconciliation and replay workflows require FinOps training and runbook adoption.\n- **Cost**: Additional RDS storage (~2 TB/mo hot, S3 archive thereafter), one new SNS topic during transition, no material DynamoDB cost change. Estimated < 8% increase in payments-platform infra cost.\n\n**Review date**: 6 months post-launch, or on any P1 payment integrity incident.",
  "cost": 0.509515,
  "finishReason": "stop"
}