Guide a Research Question down to Concepts, Dimensions, Metrics, Variables, and the Outcomes they inform — then export a detailed, programmer-ready CSV. This workspace starts empty and holds only your team's operational definitions — not the underlying evaluation data itself.
Your mappings, products, and registered data sources are persisted to a local SQLite database (/tmp/operationalization_mapper.db) and survive restarts. Bring your own data via CSV import, or load the worked example to see the pattern.
Question → Concept → Dimension → Metric → Outcome. Hover any node for details. Metric→Outcome links are many-to-many (a metric can proxy several outcomes, an outcome can be evidenced by several metrics), so the layered layout is a directed graph, not a strict tree.
These headers populate the Variables multi-select in the Add Mapping form.
No mappings yet.
Showing the core columns; click Details for derivation logic, validity/bias/provenance, and framework reference.
No rows available.
No data source headers registered yet.
Established healthcare evaluation frameworks, shipped as reference content so this is populated from the first run. Tag a Concept/Dimension/Metric against a construct here (Framework Reference field) to anchor your operational definitions to prior art instead of reinventing them — and read the "assumptions to interrogate" before trusting a framework's numbers at face value.
Consolidated Framework for Implementation Research — determinants of implementation success.
Interrogate: CFIR constructs are determinants of implementation, not outcomes themselves — don't report a CFIR construct score as if it were a clinical or operational outcome.
Damschroder et al. (2009), Implementation Science.
Classic healthcare quality model separating structural capacity, care processes, and resulting outcomes.
Interrogate: Process metrics (e.g. alerts fired, notes generated) are often reported as if they were outcomes; keep them explicitly labeled as process measures unless a validated link to a downstream outcome has been established.
Donabedian (1966), Milbank Memorial Fund Quarterly.
Grading of Recommendations Assessment, Development and Evaluation — rates certainty of evidence for an effect estimate.
Interrogate: A single observational deployment study, however large, rarely supports a 'high certainty' GRADE rating — check that the certainty label matches study design, not just sample size.
Guyatt et al. (2008), BMJ.
Nonadoption, Abandonment, Scale-up, Spread, Sustainability framework for complex health tech.
Interrogate: Easy to treat 'Technology' domain metrics (uptime, latency) as sufficient evidence of readiness while ignoring Organization/Adopter System domains, which usually predict abandonment better than technical performance.
Greenhalgh et al. (2017), J Med Internet Res.
Single-item 'likelihood to recommend' score, net of promoters minus detractors.
Interrogate: Response rates for NPS surveys are typically low and non-random; treat NPS as a directional customer-feedback signal, not a validated clinical or operational outcome, and always report response rate alongside the score.
Reichheld (2003), Harvard Business Review.
Program-evaluation framework for translating interventions into real-world impact.
Interrogate: Reach and Adoption are often conflated with raw usage counts; without a denominator of eligible-but-non-adopting users, apparent 'high adoption' can mask serious selection bias in who actually uses the product.
Glasgow, Vogt & Boles (1999), Am J Public Health.
ONC self-assessment guides for safe and effective use of health IT/EHR-adjacent systems.
Interrogate: SAFER Guides are self-assessment checklists (process presence/absence), not outcome measures — completing a checklist item is not evidence the safeguard actually functions under real load.
ONC SAFER Guides, HealthIT.gov.
10-item standardized questionnaire producing a 0-100 usability score.
Interrogate: SUS is validated for relative comparison across systems, not as an absolute pass/fail threshold; respondents are typically self-selected volunteers, which can bias scores toward more engaged/satisfied users.
Brooke (1996).
Technology Acceptance Model / Unified Theory of Acceptance and Use of Technology — predict usage intention from perceived usefulness and ease of use.
Interrogate: Self-reported intention-to-use surveys often diverge from observed usage logs; treat survey-based TAM/UTAUT scores as a distinct (and separately biased) signal from behavioral log data, not a substitute for it.
Davis (1989); Venkatesh et al. (2003).