Reality Skill Research

Reality Skill itself has to face reality.

Evidence for individual ingredients is not evidence for the integrated protocol. We separate what is independently supported, what is a working hypothesis, what Experiment #001 measures, and what would force us to revise the system.

evidence levelsfeasibility pilothuman-only baselinepublish the update history
What research already gives us

Some ingredients have meaningful evidence. The combined Reality Skill system does not yet.

Forecasting

Explicit probabilities can be trained and scored

Forecasting research supports calibration practice, teaming and systematic aggregation in appropriate domains. It does not prove every personal or business decision becomes better.

Debiasing

Some targeted interventions transfer

Specific debiasing practices can change behavior beyond the exercise, but transfer is intervention-specific and should be measured rather than assumed.

Groups

Groups can improve judgment - or destroy independence

Structured independent elicitation can preserve diverse signals; ordinary social influence can collapse the diversity that makes aggregation useful.

Verification

Source evaluation is a trainable behavior

Lateral reading and provenance checks can improve evaluation of unfamiliar information sources.

Execution

If–then planning can help convert intention into action

Explicit update triggers are preferable to hoping we will remember to revise a plan later.

Systems

Frameworks can diagnose without being universal prediction engines

Systems and constraint models are useful when they generate testable implications rather than becoming explanatory dogma.

Evidence levels

Evidence level changes how confidently a model is used.

A

Strong empirical support

Repeated empirical support for the underlying mechanism or method.

B

Supported practice

Important ingredients are supported, while this exact transfer or use may be less directly established.

C

Diagnostic framework

Useful for organizing complexity and generating interventions, but not treated as one complete predictive theory.

D

Hypothesis generator

A lens that may reveal possibilities but should not outrank stronger evidence.

The integration hypothesis

Reality Skill combines established ingredients with unproven integration choices.

Independently studied ingredients
  • Scientific testing and falsification
  • Base rates and outside-view reasoning
  • Probabilistic forecasting and calibration
  • Causal and counterfactual reasoning
  • Premortem and structured challenge
  • Source verification and lateral reading
  • Implementation intentions and feedback
  • Structured independent elicitation
Reality Skill hypotheses
  • Reality Map as a first-class input layer
  • Async-first routing before live discussion
  • Human signals with provenance
  • Scored Journal as a compounding meta-layer
  • Monthly Chartroom as an escalation + trust layer
  • Foundation Gate based on observed reliability/candor/independence
The integrated protocol has n=0 before the founding cohort. Protocol v1.2 is a testable composite hypothesis rather than a finished doctrine.
Reality Skill Experiment #001

Minsk Founding Cohort · human-only feasibility pilot.

ElementProtocol v1.2
Participants6–8 selected participants
FoundationApproximately 10–12 weeks; 6 in-person meetings
Foundation GateReliability, decision-relevant candor and independent-first behavior must be working before monthly cruise
CruiseOne in-person Chartroom every 4 weeks + limited async panels
Async capacityApproximately 3–4 Full Panels per group/month; typical member 2–4 responses/month
AINot used in the founding group decision process; future research layer only
Primary aimFind observable decision/process signal worth testing further while keeping participant burden low
Primary questions

What Experiment #001 is actually trying to learn.

01

Decision throughput

What proportion of consequential cases can be resolved without a synchronous group meeting?

02

Independent judgment

Does private-first async elicitation preserve meaningful differences before social influence?

03

Unique information

Does the panel surface relevant observations and first-hand experience the owner did not already possess?

04

Escalation quality

Do cases that reach the Chartroom actually contain more valuable disagreement, competing models or tacit information?

05

Participant burden

How much total participant time is required per resolved decision?

06

Trust + candor

Does Foundation create enough trust for decision-relevant versions of sensitive real cases?

07

Monthly sustainability

Can one in-person meeting every four weeks maintain the relationships/context needed for async cooperation?

08

Learning

Does the Scored Journal improve forecasting, review behavior and willingness to update?

09

Facilitator dependence

Can the protocol function increasingly without an unusually strong founder-facilitator?

10

AI marginal value - later

Once a stable human baseline exists, does any AI layer improve results enough to justify complexity and privacy cost?

Foundation Gate is itself a hypothesis

Transition to monthly is data-triggered.

Reliability

Working threshold: 80%+ responses without individual reminders

A practical signal that async responsibility survives without constant facilitator effort.

Decision-relevant candor

Cases contain enough real context to reason from

Not “disclose everything”; disclose enough that omitted context is unlikely to reverse the judgment.

Independence

Private-first behavior and unique contribution are visible

Genuine dissent when it exists; no requirement to manufacture disagreement.

What we track

Behavior, burden and outcomes - not enthusiasm alone.

  • Consequential decisions logged
  • Initial and final estimates
  • Cases closed async
  • Cases escalated to Chartroom
  • Unique information added
  • Time spent per resolved decision
  • Actions, update triggers and review dates
  • Resolved outcomes
  • Forecast calibration where applicable
  • Async response reliability
  • Decision-relevant candor
  • Relationship continuity
  • Emergency Mini-Room frequency
What would make us change our mind?

A self-correcting framework should expose its own failure conditions.

  • Participants enjoy the process but consequential decisions do not materially change.
  • Tools work only inside meetings and do not transfer into ordinary work or life.
  • Async panels become sanitized, low-response or socially conformist.
  • Group discussion consistently reduces independence or produces worse judgments.
  • The benefit depends heavily on one unusually strong facilitator.
  • The protocol creates more cognitive overhead than decision value.
  • Monthly cadence is too sparse to sustain trust and candor.
  • Repeated cohorts fail to produce a measurable signal worth pursuing.
Research roadmap

Earn stronger claims gradually.

Stage 01

Feasibility

Can people actually use the protocol on real decisions, and can we measure it?

Stage 02

Replication

Run more cohorts and separate robust patterns from founder/group effects.

Stage 03

Comparative tests

Compare structured solo, async independent panels, ordinary discussion, structured room and later AI additions where appropriate.

Stage 04

Publish updates

Report useful results, null results and protocol changes so Reality Skill itself has an update history.

Working bibliography

Research traditions currently informing the protocol.

Forecasting & aggregation

Mellers and colleagues; Good Judgment work; structured expert elicitation and IDEA/Delphi research.

Social influence & collective judgment

Wisdom of crowds, hidden profiles, information pooling, conformity and collective intelligence.

Debiasing & decision preparation

Considering the opposite, premortem-style challenge, implementation intentions, calibration and outcome review.

External cognition & systems

Cognitive offloading, external representations, systems thinking, constraints, feedback, bottlenecks and value of information.

Keep the existing detailed source links below this section if you want to preserve the current bibliography verbatim.

Important limitation: 6–8 participants cannot scientifically validate Reality Skill. Experiment #001 is a feasibility and measurement pilot intended to expose signal, failure modes and better hypotheses.
Research rule

Research tells us what might work. Practice tells us whether it works here.

Delivery format · secondary research question

The primary hypothesis is decision quality, not meeting format.

In-person participation in Minsk is preferred when available because it offers richer human contact. Online participation exists for people who cannot reliably attend locally. We will still compare decision value, candor, trust and burden across delivery modes, but format remains a secondary implementation question rather than the identity of Reality Skill.