Whether a valid human-gold case and Twin result reach the same decision outcome.
Can AI reproduce how an enterprise architect actually decides?
This preliminary experiment tests more than whether an AI reaches the same final answer as an architect. It compares observable decision artefacts: evidence, principles, constraints, uncertainty, counterfactual behaviour and decision outcomes.
Decision fidelity is not a final-answer scoreboard.
A matching recommendation is a weak criterion. Enterprise architecture decisions are shaped by evidence, constraints, principles, unknowns, policy boundaries and what changes when assumptions move.
What would count as reproducing a decision?
The public graph exposes evidence references, but precision and recall are not computed as scores in this release.
Principles are visible in the decision engine, but overlap is not yet scored independently.
Unavailable human gold and private context remain UNKNOWN or N/A instead of being counted as failures.
The same deterministic engine reruns under changed constraints and reports deltas.
The structured critic exposes weakened assumptions without implying an external model critic.
Evolution deltas exist, but temporal scoring is not computed here.
2 / 2 observed agreement
N = 2; preliminary public calibration. This sample is too small to support general claims about architectural decision fidelity.
Human-gold means a decision for which a defensible public authored Prasad decision record exists. The benchmark is constrained by available public human-authority records. The Twin must never manufacture the human answer after the fact.
MATCH: Governed SAP AI needs evidence, authority boundaries and human control.
MATCH: Use read-only proof paths before execution.
UNKNOWN
External model critic unavailable
The experiment sequence uses existing R5 data.
The page does not build another reasoning engine. It points to the existing Digital Twin workbench for the runnable decision, counterfactual, fidelity, critic and receipt views.
Counterfactual: BTP unavailable
- Baseline constraint
- SAP semantic authority required, BTP unavailable
- Baseline decision
- Capella
- Changed constraint
- BTP unavailable
- Re-evaluated decision
- Capella
- Decision delta
- The strongest evidence-backed option remained stable under the changed inputs.
The value is the recalculation, not forcing the recommendation to move.
Open this scenario in the Digital TwinEvidence-first SAP AI
Matching the final recommendation does not demonstrate equivalent reasoning.
Compare in the Digital TwinChallenge the decision without inventing a model critic.
- Current decision
- Challenge the strongest recommendation.
- Challenge
- Test whether public evidence, policy and unknown-state rules weaken the recommendation.
- Finding
- The active subject includes research/prototype evidence and cannot be promoted to SUPPORTED from that signal alone.
- Effect
- DECISION UNCHANGED; risk remains visible
External model critic unavailable in R5/R5.1; the visible result is the structured deterministic critic.
Challenge in the Digital TwinWhat this does not prove.
- Very small human-gold sample.
- Public evidence is incomplete.
- Private client context is intentionally excluded.
- SAP live state may be unavailable.
- Fidelity is measured only where authoritative human records exist.
- Generated reasoning is not evidence.
- The same final answer can arise from different decision paths.
- No claim of human cognition replication.
- No claim that the Twin replaces human architectural authority.
Public artefacts behind this page.
- Decision receipts
- receipt:twin-boundary, receipt:r5-living-state, assessment-receipt:r5-boundary
- Receipt claim
- The public Prasad Digital Twin is read-only, public-safe, evidence-backed and human-accountable.
- Public MCP
- READ-ONLY
- Write tools
- 0
