Search across Architecture IntelligenceTry engineering explorations, SAP Clean Core, evidence, integration or AI governance.

Workspace

Saved on this device only — nothing here is synced or uploaded.

Digital Twin Fidelity

Can AI reproduce how an enterprise architect actually decides?

This preliminary experiment tests more than whether an AI reaches the same final answer as an architect. It compares observable decision artefacts: evidence, principles, constraints, uncertainty, counterfactual behaviour and decision outcomes.

ExperimentalStatus
2Human-gold cases
prasad-digital-twin-r5-r51-living-twin-convergenceTwin snapshot
11Evidence records
0Write tools
01 Research question

Decision fidelity is not a final-answer scoreboard.

A matching recommendation is a weak criterion. Enterprise architecture decisions are shaped by evidence, constraints, principles, unknowns, policy boundaries and what changes when assumptions move.

What would count as reproducing a decision?

Decision AgreementMeasured now

Whether a valid human-gold case and Twin result reach the same decision outcome.

Evidence Precision / RecallPLANNED / NOT YET MEASURED

The public graph exposes evidence references, but precision and recall are not computed as scores in this release.

Principle OverlapPLANNED / NOT YET MEASURED

Principles are visible in the decision engine, but overlap is not yet scored independently.

UNKNOWN CalibrationObserved qualitatively

Unavailable human gold and private context remain UNKNOWN or N/A instead of being counted as failures.

Counterfactual ConsistencyObserved qualitatively

The same deterministic engine reruns under changed constraints and reports deltas.

Challenge RecoveryObserved qualitatively

The structured critic exposes weakened assumptions without implying an external model critic.

Temporal ConsistencyPLANNED / NOT YET MEASURED

Evolution deltas exist, but temporal scoring is not computed here.

02 Current public calibration

2 / 2 observed agreement

N = 2; preliminary public calibration. This sample is too small to support general claims about architectural decision fidelity.

Human-gold means a decision for which a defensible public authored Prasad decision record exists. The benchmark is constrained by available public human-authority records. The Twin must never manufacture the human answer after the fact.

Human-gold availableEvidence-first SAP AI

MATCH: Governed SAP AI needs evidence, authority boundaries and human control.

Human-gold availableRead-only before write

MATCH: Use read-only proof paths before execution.

Human-gold unavailablePrivate client facts

UNKNOWN

Human-gold unavailableModel critic without provider

External model critic unavailable

03 Live / current evidence

The experiment sequence uses existing R5 data.

EvidenceTwin statePrinciples + constraintsDecisionCounterfactualHuman comparisonChallengeReceipt

The page does not build another reasoning engine. It points to the existing Digital Twin workbench for the runnable decision, counterfactual, fidelity, critic and receipt views.

04 Counterfactual test

Counterfactual: BTP unavailable

Baseline constraint
SAP semantic authority required, BTP unavailable
Baseline decision
Capella
Changed constraint
BTP unavailable
Re-evaluated decision
Capella
Decision delta
The strongest evidence-backed option remained stable under the changed inputs.

The value is the recalculation, not forcing the recommendation to move.

Open this scenario in the Digital Twin
05 Compare with Prasad

Evidence-first SAP AI

DimensionHumanTwinState
Decisionpublication:evolving-sap-development-architectGoverned SAP AI needs evidence, authority boundaries and human control.AGREEMENT
Evidencepublication:evolving-sap-development-architectCanonical public graph referenceAGREEMENT
PrinciplesEvidence authority and human controlEvidence before recommendation; read-only firstAGREEMENT
UnknownsPrivate context not presentUNKNOWN remains explicitAGREEMENT

Matching the final recommendation does not demonstrate equivalent reasoning.

Compare in the Digital Twin
06 Structured critic

Challenge the decision without inventing a model critic.

Current decision
Challenge the strongest recommendation.
Challenge
Test whether public evidence, policy and unknown-state rules weaken the recommendation.
Finding
The active subject includes research/prototype evidence and cannot be promoted to SUPPORTED from that signal alone.
Effect
DECISION UNCHANGED; risk remains visible

External model critic unavailable in R5/R5.1; the visible result is the structured deterministic critic.

Challenge in the Digital Twin
07 Limitations

What this does not prove.

  • Very small human-gold sample.
  • Public evidence is incomplete.
  • Private client context is intentionally excluded.
  • SAP live state may be unavailable.
  • Fidelity is measured only where authoritative human records exist.
  • Generated reasoning is not evidence.
  • The same final answer can arise from different decision paths.
  • No claim of human cognition replication.
  • No claim that the Twin replaces human architectural authority.
08 Inspect / reproduce

Public artefacts behind this page.

Decision receipts
receipt:twin-boundary, receipt:r5-living-state, assessment-receipt:r5-boundary
Receipt claim
The public Prasad Digital Twin is read-only, public-safe, evidence-backed and human-accountable.
Public MCP
READ-ONLY
Write tools
0