How Much Evidence Is Enough?
How to set evidence thresholds before hope, fear or momentum make the choice—and decide what the next commitment genuinely requires.
No important decision arrives with complete evidence.
There is always something still unknown.
A customer may behave differently after launch.
A partner may change direction.
A community may respond in ways the organisation did not anticipate.
A technology may perform well during testing and fail under real operating conditions.
Costs may increase.
Demand may weaken.
A funding decision may take longer than expected.
A competitor may enter the market.
The wider context may change.
This creates one of the most difficult questions in developing an idea:
How much evidence is enough?
Too little evidence, and an organisation may commit money, time, reputation and trust to a proposition built largely on hope.
Too much research, and the opportunity may disappear while the organisation continues investigating questions that can only be answered through action.
The objective cannot be certainty.
Certainty is rarely available before a consequential decision.
The more useful objective is justified confidence.
Do we have enough relevant evidence to make this particular commitment, at this particular scale, under these particular conditions?
That question is different from:
Have we proved the idea will succeed?
Success cannot usually be proved in advance.
But the basis for the next decision can be made visible.
That is the role of a decision threshold.
A decision threshold defines what an organisation needs to see before it will:
proceed;
revise;
pause;
or stop.
It establishes the level and type of evidence required before greater commitment becomes justified.
Most importantly, it should be defined before enthusiasm, fear, attachment or momentum begin interpreting every result in their own favour.
Because once people become emotionally or financially invested in an idea, “enough evidence” has a habit of changing.
Evidence Cannot Remove Every Uncertainty
The desire for more evidence is understandable.
Evidence can reveal whether people experience the problem we believe they experience.
It can test whether they understand the proposition.
It can indicate whether they are willing to participate or pay.
It can expose operational weaknesses.
It can challenge cultural assumptions.
It can identify barriers.
It can show where a project needs revision.
But evidence does not convert the future into a known fact.
A successful pilot does not guarantee a successful national launch.
Ten committed participants do not prove that ten thousand will participate.
A technically successful prototype does not establish that the organisation can operate the service reliably.
A signed partnership does not guarantee that the relationship will remain productive.
Positive community consultation does not prove that participation will be sustained.
Evidence reduces uncertainty.
It does not abolish it.
This distinction matters because organisations often move between two unhelpful extremes.
The first says:
“We cannot know everything, so we may as well trust our instincts.”
The second says:
“We cannot act until we are certain.”
Neither position is adequate.
The first can disguise weak reasoning as courage.
The second can disguise avoidance as rigour.
A more intelligent position is:
We will never know everything. What do we need to know before taking the next proportionate step?
That is the question a decision threshold is designed to answer.
“Enough” Depends Upon What You Are About to Do
There is no universal quantity of evidence that makes an idea safe.
The appropriate threshold depends upon the commitment being considered.
A conversation requires very little evidence.
A prototype requires more.
A paid pilot requires more again.
A limited launch may require evidence about demand, delivery, price and customer experience.
A major public commitment, long-term contract or substantial investment should normally require stronger and more varied evidence.
The principle is simple:
The greater the commitment, the stronger the evidence should be.
But commitment is not measured by money alone.
It can also involve:
time;
reputation;
relationships;
public trust;
staff capacity;
community expectations;
personal wellbeing;
data;
legal responsibility;
environmental consequences;
or the opportunity cost of not doing something else.
A £2,000 project may be reversible for one organisation and existential for another.
A small community event may have a modest budget but significant consequences if it damages trust.
An experimental AI tool may appear financially inexpensive while introducing serious risks around privacy, reliability or discriminatory outcomes.
Evidence thresholds must therefore be proportionate to the real consequences of being wrong.
Five Factors That Should Raise the Threshold
Before deciding how much evidence is enough, examine five dimensions.
1. Scale of Commitment
How much money, time, attention and capacity will the next step require?
A larger commitment usually demands stronger evidence.
2. Consequence of Error
What happens if the assumption proves false?
Will the organisation make a small adjustment?
Or could people lose money, access, trust, employment, opportunity or safety?
The severity of the possible consequence matters.
3. Reversibility
How easily can the decision be changed?
A temporary landing page can be revised quickly.
A long property lease cannot.
A pilot service can be stopped.
A public promise to a community may be much harder to withdraw without consequence.
Irreversible or difficult-to-reverse decisions deserve higher thresholds.
4. Who Is Affected
Does the risk fall only on the decision-maker?
Or are customers, staff, partners, artists, participants, communities or vulnerable people carrying part of it?
The more risk is transferred to other people, the stronger the responsibility to investigate it properly.
5. Cost of Delay
What could be lost by waiting?
More research is not free.
Delay consumes time, money and attention.
It may allow a funding window to close, reduce momentum, postpone learning or leave an important need unaddressed.
A good threshold considers both the cost of acting too early and the cost of acting too late.
Together, these factors establish the seriousness of the decision.
A Threshold Is Not a Feeling
Without a threshold, decisions often depend upon emotional interpretation.
An optimistic founder sees three positive comments and concludes that demand has been validated.
A fearful founder sees one objection and concludes that the idea will fail.
A team under deadline pressure decides the available evidence is sufficient because the launch date has already been announced.
An organisation that has invested heavily continues because stopping would require admitting that earlier assumptions were weak.
In each case, the evidence has not necessarily changed.
The standard applied to it has.
A threshold protects the decision from this drift.
For example:
“We will proceed to a paid pilot if at least twelve people from the intended audience commit at the proposed price, at least ten complete the experience, and the pilot can be delivered within the agreed cost and capacity limits.”
Or:
“We will not proceed to public launch until the safeguarding process has been independently reviewed, the service has performed reliably during a controlled test and responsibility for human oversight has been assigned.”
These thresholds are imperfect.
But they make the reasoning visible.
The decision is no longer based solely on whether the team feels encouraged.
It is connected to conditions defined in advance.
Set the Threshold Before You See the Result
This is one of the most important principles in the process.
Define what would count before the evidence arrives.
Suppose an artist launches a pre-order campaign for a limited-edition print.
Before the campaign, they decide:
Proceed with the full edition if at least twenty-five orders are placed at the intended price within thirty days.
Revise the offer if interest is high but fewer than fifteen people purchase.
Pause if the audience engages but production costs make the offer unviable.
Stop the edition if there is little qualified interest after the offer has reached the intended audience.
Now the results can be interpreted against a prior standard.
Without that standard, the artist may receive six orders and say:
Six people paid. That proves demand.
Or:
Only six people paid. The idea is a failure.
Neither conclusion necessarily follows.
Six orders might justify a smaller edition.
They might indicate that the price is acceptable but the audience is too limited.
They might suggest that the work resonates with a particular group.
They might reveal that the campaign did not reach enough relevant people.
The threshold does not eliminate interpretation.
It prevents interpretation from becoming completely untethered.
Moving the Goalposts
Once results arrive, teams often begin renegotiating what they originally said would matter.
A pilot needed twenty participants.
Only eight joined.
The team now argues that eight highly engaged participants are more valuable than twenty less engaged ones.
That may be true.
But it may also be a rationalisation.
The right response is not necessarily to abandon the project.
It is to document the change honestly.
For example:
Original threshold: twenty paid participants.
Actual result: eight paid participants, six of whom completed and requested a follow-on offer.
Interpretation: the scale of demand was lower than expected, but repeat interest suggests possible value within a narrower audience.
Decision: revise the proposition and test a smaller specialist version rather than treating the original demand assumption as validated.
This preserves learning.
Moving the goalposts hides it.
A changed threshold should therefore be treated as a new decision, not quietly rewritten as though it had always been the standard.
Evidence Strength Is Not the Same as Evidence Volume
More information does not automatically create better evidence.
An organisation could collect:
2,000 social-media likes;
500 page views;
100 positive comments;
50 survey responses;
and ten enthusiastic messages.
This may create a strong impression of momentum.
But if the decision is whether people will pay £300 for a service, none of those signals directly answers the question.
One paid booking may be more relevant than hundreds of likes.
Five completed pilots may reveal more about delivery than an extensive opinion survey.
A detailed conversation with an affected community may expose risks that a large generic dataset cannot see.
Evidence should therefore be assessed by quality, not merely quantity.
Several dimensions matter.
Strength
How directly does the evidence demonstrate the behaviour, capacity or condition being claimed?
A statement of interest may be weaker than a deposit.
A deposit may be weaker than a completed purchase.
A completed purchase may be weaker than repeated purchase when the proposition depends upon retention.
The closer the evidence is to the real commitment, the stronger it may be.
Relevance
Does the evidence answer the actual question?
Website visits are relevant to attention.
They are not direct evidence of willingness to pay.
Positive feedback is relevant to appeal.
It is not direct evidence of delivery capacity.
A successful technical demonstration is relevant to feasibility.
It may not establish customer acceptance, organisational readiness or cultural appropriateness.
Evidence can be strong in itself and still irrelevant to the decision being made.
Convergence
Do different sources point in a similar direction?
Suppose interviews reveal a repeated problem.
Search behaviour suggests people actively seek help with it.
A paid pilot attracts participants.
Several participants return.
Comparable external evidence indicates similar demand elsewhere.
No single source is conclusive.
Together, they create convergence.
Confidence should increase when independent forms of evidence support the same interpretation.
This is why the Cultural Intelligence Simulation Lab uses an Evidence Ledger rather than relying upon one attractive statistic or quotation.
Behavioural Evidence and Stated Evidence
What people say matters.
Interviews, conversations and consultation can reveal:
meaning;
motivation;
trust;
concerns;
language;
cultural context;
and how people interpret an idea.
But stated intention and actual behaviour are not identical.
A person can sincerely say:
“I would definitely buy that.”
Then not buy it.
A community member can support an idea during consultation but be unable to attend because timing, transport, cost or caring responsibilities create barriers.
A partner can express strong enthusiasm but be unable to commit staff or money.
Behavioural evidence shows what people actually do under particular conditions.
Stated evidence helps explain why.
Strong decision-making often needs both.
Behaviour without interpretation can be misunderstood.
Interpretation without behaviour can overstate commitment.
Representativeness
Whose experience does the evidence describe?
A founder may test an idea with friends and receive enthusiastic feedback.
But friends may be unusually supportive.
A cultural organisation may consult its existing audience.
But the programme may be intended for people who do not currently engage.
An online survey may attract people already interested in the subject.
A pilot may exclude those facing the greatest access barriers.
Evidence does not need to represent everybody.
But the organisation should understand whom it represents—and whom it does not.
Ten carefully selected interviews can be more decision-useful than 500 responses from the wrong population.
The issue is not simply sample size.
It is whether the evidence comes from people relevant to the claim.
Recency
When was the evidence collected?
Customer behaviour changes.
Costs change.
Platforms change.
Community relationships change.
Technology develops.
Funding priorities shift.
A result that was highly relevant two years ago may no longer justify the same confidence.
This does not make older evidence useless.
It means its current relevance should be examined rather than assumed.
Contradiction
What evidence points in the opposite direction?
Teams often collect supportive evidence and treat conflicting evidence as noise.
But contradiction can be valuable.
Suppose customers like the proposition but refuse the price.
Suppose participants value the programme but cannot attend at the proposed time.
Suppose the technology works but staff do not trust it.
Suppose partners support the mission but will not commit resources.
The contradiction may reveal where the idea needs revision.
Evidence that complicates the story is not automatically bad evidence.
It may be the evidence doing the most useful work.
Unknowns
What important questions remain unanswered?
A decision can sometimes proceed despite major uncertainty.
But the uncertainty should be visible.
An Evidence Ledger should not create a false impression that every box has been completed.
It should distinguish:
what is supported;
what is reasonably inferred;
what remains assumed;
what has been contradicted;
and what is still unknown.
The unknown category is not a failure.
It is part of intellectual honesty.
When Evidence Conflicts
Real evidence rarely arrives in a perfectly coherent form.
A founder may hear strong enthusiasm during interviews but see weak conversion during a test.
A community programme may attract substantial attendance while participants report that the experience felt tokenistic.
An AI system may save staff time while producing outputs that require significant correction.
A product may receive positive reviews but generate insufficient margin.
What should happen?
Begin by resisting the urge to average everything into one conclusion.
Ask which part of the proposition each source addresses.
For example:
High attendance may support the claim that people are interested.
Critical participant feedback may challenge the claim that the programme is culturally or practically well designed.
Both can be true.
Similarly:
An AI tool may demonstrate technical usefulness.
The correction burden may challenge the claim that it produces an operational saving.
Again, both can be true.
Conflicting evidence often indicates that the proposition is not simply good or bad.
Different parts are performing differently.
That may lead towards revision rather than a binary verdict.
False Precision
Decision thresholds can improve discipline.
They can also create the appearance of certainty where none exists.
A team might score an idea:
Demand: 7.4
Feasibility: 6.8
Cultural fit: 8.1
Overall readiness: 74%
The numbers look authoritative.
But what do they actually represent?
If the scoring system is built upon subjective estimates, the decimal points do not make it objective.
Measurement can be useful.
False precision is not.
Thresholds should be specific enough to guide a decision but honest about the quality of the underlying evidence.
Sometimes the most accurate conclusion is:
The available evidence supports a limited next step, but not a larger commitment.
That sentence may be more useful than a sophisticated-looking score.
Endless Research Is Also a Decision
Evidence-led practice can become an avoidance mechanism.
Another interview.
Another survey.
Another market report.
Another workshop.
Another month of analysis.
Eventually, the organisation is no longer reducing meaningful uncertainty.
It is postponing exposure to reality.
Some questions cannot be answered through further desk research.
Will customers pay?
Offer them the opportunity.
Can the organisation deliver?
Run a pilot.
Will the venue work for the intended participants?
Test the experience with them.
Can the team operate the system safely?
Simulate the workflow.
Research should continue while it is capable of changing the decision.
When the remaining uncertainty can only be reduced through action, the next responsible step may be a bounded experiment.
This is where Article Three’s smallest useful test becomes essential.
The objective is not to finish learning before acting.
It is to choose an action designed to continue the learning.
Validation Theatre
There is a difference between testing an idea and staging a performance in which the idea always wins.
Validation theatre happens when:
questions are designed to invite praise;
only supportive participants are consulted;
weak signals are presented as strong proof;
contradictory evidence is excluded;
the price is never tested;
delivery capacity is ignored;
or the criteria for success are rewritten after the results arrive.
The process looks rigorous.
The outcome was never genuinely open.
A useful test must contain the possibility that the organisation learns something it did not want to hear.
If no result could cause the idea to be revised, paused or stopped, the exercise is not testing the decision.
It is collecting reassurance.
Different Decisions Need Different Evidence
The phrase “we need more evidence” is incomplete.
Evidence of what?
For which decision?
At what level of commitment?
A practical evidence threshold should match the type of claim being made.
Commercial Demand
If the claim is:
“Customers will pay for this,”
use evidence that moves beyond general approval.
Relevant signals might include:
paid pilots;
pre-orders;
deposits;
completed purchases;
repeat purchases;
retention;
credible letters of intent;
or real sales conversations at the proposed price.
Likes, compliments and expressions of interest may be useful early signals.
They should not be treated as equivalent to purchasing behaviour.
The threshold should also consider whether the economics work.
Demand at a loss-making price does not automatically validate the business model.
Community Participation and Trust
If the proposition depends upon community participation, attendance alone may be insufficient.
Ask:
Did the people the programme intended to involve actually participate?
Who did not?
What barriers remained?
Did people feel listened to?
Did they have meaningful influence?
Did the experience strengthen or weaken trust?
Would they participate again?
Did the organisation describe the relationship honestly?
A full room can demonstrate attendance.
It does not automatically demonstrate trust, inclusion, shared ownership or long-term relevance.
Evidence thresholds should reflect the real claim.
Funding and Partnership Commitments
A positive conversation is not a secured partnership.
A promising funding opportunity is not an award.
A verbal expression of interest is not committed capacity.
Evidence may strengthen through stages:
initial interest;
follow-up meeting;
agreed role;
written confirmation;
allocated staff time;
committed money or resources;
signed agreement.
The appropriate threshold depends upon how heavily the project relies on the partner or funder.
If the project cannot operate without them, informal enthusiasm should not be treated as infrastructure.
Organisational Capacity
If the claim is:
“We can deliver this,”
evidence should include more than technical possibility.
Consider:
staff time;
skills;
cash flow;
systems;
supplier reliability;
quality control;
customer support;
leadership attention;
contingency capacity;
and the effect on existing work.
A pilot that succeeds only because the founder works unsustainable hours may validate customer interest while disproving the delivery model.
Capacity evidence must reflect how the service would operate under realistic conditions.
Artificial Intelligence and Technology
If a proposition depends upon AI or another technology, the evidence standard should extend beyond whether the tool can produce an impressive demonstration.
Ask:
Does it work reliably across realistic cases?
How often does it fail?
Can failures be detected?
What level of human review is required?
Does it genuinely save time after checking and correction?
What data does it use?
What privacy, security or consent issues arise?
Who is accountable for consequential decisions?
Does performance differ across people, languages or contexts?
Can the organisation operate if the system becomes unavailable or changes?
A successful demonstration may justify a controlled test.
It should not automatically justify deployment into high-consequence decisions.
The greater the potential effect on people, the stronger the threshold for reliability, oversight and accountability should become.
What Is Enough for the Next Stage?
A useful way to avoid both premature commitment and endless research is to match evidence to the next stage.
Enough for a Conversation
You need a plausible observation, a clearly expressed problem or an informed question.
The purpose is exploration.
The risk is low.
Enough for a Prototype
You need a credible problem worth investigating, a testable idea and sufficient reason to believe that building a small representation could answer something important.
The prototype does not need to prove the business.
It needs to generate learning.
Enough for a Pilot
You need evidence that the problem matters to a relevant audience, that the proposed approach is sufficiently credible to test, and that the pilot can be delivered responsibly.
Where people are affected, appropriate safeguards, consent and oversight should already exist.
Enough for a Limited Launch
You need stronger evidence of demand or participation, operational feasibility, appropriate pricing or funding, delivery capacity and a way to monitor outcomes.
The organisation should know what would trigger revision or withdrawal.
Enough for a Major Commitment
You need converging evidence across the proposition’s major dependencies.
This might include demonstrated demand, viable economics, delivery capability, committed partners, cultural and ethical consideration, governance, realistic scenarios and contingency plans.
Unknowns may remain.
But the organisation should understand which unknowns it is consciously accepting.
This staged approach recognises that evidence should grow with commitment.
It prevents a small encouraging signal from being used to justify a disproportionately large decision.
Practical Example One: The Creative Founder
A creative founder plans to produce 500 units of a premium object.
The original belief is:
Customers will value the craftsmanship and pay £180.
Available evidence includes:
positive social-media responses;
encouragement from existing followers;
ten customer interviews;
and strong reactions to prototype images.
Useful, but incomplete.
The major commitment is the production run.
The load-bearing assumptions involve:
willingness to pay;
production quality;
unit economics;
and the ability to reach enough relevant buyers.
A proportionate threshold might be:
Proceed to a limited production run if at least thirty customers place deposits at £180, the manufacturer produces three samples meeting the quality standard, and the final unit economics preserve the required margin.
Revise if customers place deposits only at a lower price or if production costs make the current design unviable.
Pause if the prototype attracts interest but quality cannot be delivered consistently.
Stop the current version if qualified customers repeatedly reject both the price and the core proposition.
The threshold does not predict success.
It determines what would justify the next commitment.
Practical Example Two: The Cultural or Community Programme
A cultural organisation wants to develop a programme with a local community.
The organisation has evidence that:
the subject is relevant;
several community leaders are interested;
and previous events attracted reasonable attendance.
But the proposed programme depends upon sustained participation from people who have not yet been involved.
A responsible threshold might include:
participation from the intended groups during programme design;
evidence that the format and venue are accessible;
clarity about who holds decision-making power;
written commitments from delivery partners;
a realistic staffing plan;
and confirmation that participants consider the proposition worthwhile.
Proceed to a pilot if those conditions are met.
Revise if people support the underlying purpose but reject the proposed format.
Pause if trust-building requires more time.
Stop or redesign the programme if the organisation cannot involve the people whose participation is being used to justify it.
Here, “enough evidence” is not simply a participant number.
It includes the quality of the relationship.
Practical Example Three: The AI-Enabled Small Business
A small business wants to introduce an AI customer-support assistant.
The demonstration looks impressive.
It responds quickly and handles common questions.
But the real decision is whether to allow it to interact with customers.
A staged threshold might be:
Proceed to internal testing if the system can answer a defined set of low-risk questions using approved information.
Proceed to controlled customer use if responses meet an agreed accuracy standard, errors can be detected, customers can reach a human easily, data handling is appropriate and staff know who is responsible for oversight.
Revise if the system creates time savings only by transferring correction work elsewhere.
Pause if privacy, security or accountability remain unclear.
Stop the deployment if the system repeatedly produces harmful, misleading or discriminatory outputs that cannot be reliably controlled.
The impressive demonstration was evidence of possibility.
It was not yet evidence of readiness.
Whose Evidence Counts?
Evidence is never collected outside power.
Institutions often privilege certain forms of knowledge.
Formal reports may be treated as authoritative.
Lived experience may be treated as anecdotal.
Quantitative data may be considered objective.
Community testimony may be asked to prove itself repeatedly.
Technical expertise may receive immediate credibility.
Local knowledge may be consulted only after the main decisions have already been made.
A culturally intelligent threshold asks:
Who defined what counts as evidence?
Whose experience is considered credible?
Who was able to participate in the research?
Who was excluded by the method?
Whose risk is being measured?
Whose risk is being ignored?
A dataset can be methodologically tidy while missing the people most affected by the decision.
That is not simply a research limitation.
It can become a strategic and ethical failure.
Who Bears the Risk?
The acceptable threshold should rise when the benefits and risks are distributed unevenly.
Suppose an organisation gains efficiency from automating a process, but customers bear the consequences when the system fails.
Or a cultural institution gains funding and reputation from a community programme while participants carry the emotional labour.
Or a founder gains the upside from rapid growth while freelance workers absorb unstable conditions.
The decision-maker may consider the risk manageable because someone else experiences the harm.
A responsible threshold asks:
Who benefits if this works?
Who pays if it fails?
That answer should influence how much evidence is required.
Who Is Missing?
Absence matters.
A service may test well among people who can easily access it.
What about those who cannot?
A programme may receive positive feedback from existing participants.
What about people who never entered the building?
An AI system may perform well in standard cases.
What about less common languages, communication styles or accessibility needs?
A threshold should not demand impossible universal representation.
But it should recognise when the evidence systematically excludes people relevant to the claim.
Sometimes the most important evidence is the discovery that the research has not yet reached the people whose experience matters most.
The CIS Decision Threshold Framework
The Cultural Intelligence Simulation Lab can structure decision thresholds through eight connected questions.
1. Decision
What exact decision are we trying to make?
Not:
Is this a good idea?
But:
Should we fund a pilot?
Should we produce the first edition?
Should we enter the partnership?
Should we launch to customers?
2. Commitment
What resources, relationships and responsibilities will the decision create?
Include money, time, reputation, trust, capacity and opportunity cost.
3. Critical Claims
What must be true for this commitment to be justified?
These are the load-bearing assumptions identified through the Evidence Ledger and Red Team challenge.
4. Required Evidence
What type of evidence would meaningfully support or challenge each claim?
Match the evidence to the question.
5. Thresholds
What result would lead to:
Proceed?
Revise?
Pause?
Stop?
Set these conditions before interpreting the findings.
6. Cultural and Human Consequences
Who is affected?
Whose evidence counts?
Who carries the risk?
Who may be missing?
What histories, identities, relationships or power structures shape the decision?
7. Remaining Uncertainty
What will still be unknown even if the threshold is reached?
Can that uncertainty be monitored, reduced through the smallest useful test or managed through staged commitment?
8. Review Point
When will the decision be reconsidered?
What early warning signals should trigger review?
A threshold should not be treated as permanent permission.
Conditions change.
Evidence should continue developing after the decision.
From Evidence Ledger to Decision Report
The threshold does not sit alone.
It connects the previous stages of the Before You Commit process.
The Evidence Ledger separates:
what is known;
what is inferred;
what is assumed;
what has been contradicted;
and what remains unknown.
The load-bearing assumption analysis identifies which uncertainties matter most.
The smallest useful test generates evidence efficiently.
Scenario thinking examines how the proposition behaves under different futures.
The pre-mortem imagines failure and reveals risks optimism may overlook.
The decision threshold then asks:
Given everything we now understand, what would justify the next commitment?
The CIS Decision Report can bring these elements together.
It does not simply say:
The idea scored well.
It makes the reasoning visible:
the evidence available;
the standard applied;
the uncertainties accepted;
the cultural considerations;
the conditions of the decision;
and what should happen next.
That leads towards one of four directions.
Proceed
The threshold has been met strongly enough to justify the next proportionate commitment.
Proceed does not mean:
Success is guaranteed.
It means:
The available evidence provides a reasonable basis for this next step.
The scale of that step should still match the strength of the evidence.
Revise
Part of the proposition remains credible, but something material needs to change.
Perhaps:
the audience;
price;
format;
language;
delivery model;
timescale;
partnership;
technology;
or level of ambition.
Revision can be a sign that the evidence process is working.
The objective is not to protect the original form of the idea.
It is to discover what the opportunity deserves next.
Pause
The threshold has not been met because an important uncertainty remains unresolved.
A pause should identify:
what is missing;
why it matters;
how it could be investigated;
and when the decision will be revisited.
Without those conditions, pause can become indefinite drift.
With them, it becomes an active strategic choice.
Stop
The evidence suggests that further commitment to the current proposition is not justified.
Perhaps demand is insufficient.
The economics do not work.
The delivery model is unsustainable.
The cultural or ethical consequences are unacceptable.
A critical partner is unavailable.
The technology cannot be governed responsibly.
Stopping can protect resources, relationships and future possibility.
It may also reveal where the valuable part of the idea actually lies.
A stopped proposition does not always mean the underlying insight was wrong.
Sometimes it means the current expression of it was.
The Threshold Should Match the Courage Required
It takes courage to proceed without certainty.
It also takes courage to revise an idea people have become attached to.
To pause when momentum is pushing forward.
To stop when time and identity have already been invested.
Decision thresholds do not remove the emotional difficulty.
They give the difficulty structure.
They allow a team to say:
We agreed what would matter.
This is what the evidence now shows.
This is where the threshold was met.
This is where it was not.
This is the uncertainty we are prepared to carry.
And this is the commitment that the evidence currently justifies.
That is a stronger position than pretending the future has been proved.
Cultural Intelligence Studio Perspective
Cultural Intelligence Studio does not treat evidence as a machine that produces certainty.
Evidence does not speak entirely for itself.
It has to be interpreted.
Its limitations have to be understood.
Its sources have to be examined.
Its cultural context matters.
Its absences matter.
Its contradictions matter.
And the consequences of acting upon it are rarely distributed equally.
The purpose of a decision threshold is therefore not to reduce judgement to a number.
It is to improve judgement by making the standard visible before the decision becomes emotionally difficult.
The threshold should be:
specific enough to guide action;
proportionate to the commitment;
relevant to the critical claims;
honest about uncertainty;
sensitive to cultural context;
and capable of producing more than one outcome.
If every possible result leads to proceed, there is no threshold.
There is only permission.
Conclusion: Enough Evidence for What Comes Next
The question:
“How much evidence is enough?”
cannot be answered in the abstract.
Enough for a conversation is not enough for a contract.
Enough for a prototype is not enough for a launch.
Enough for a low-risk experiment is not enough for a decision that could materially affect customers, staff, communities or public trust.
The right threshold depends upon:
the commitment;
the consequence of error;
the reversibility of the decision;
the people affected;
and the cost of waiting.
It also depends upon the quality of the evidence.
Is it strong?
Relevant?
Recent?
Representative of the people connected to the claim?
Supported by behaviour as well as stated intention?
Corroborated by different sources?
Honest about contradiction and uncertainty?
The objective is not to prove that nothing can go wrong.
That is impossible.
The objective is to know why the next step is justified—and what would cause the organisation to think again.
Set the threshold before the results arrive.
Distinguish a useful signal from a convenient one.
Do not let weak evidence carry a stronger claim than it can support.
Do not demand certainty where a reversible experiment would create better learning.
Do not continue researching when only action can answer the question.
Do not exclude the people who will bear the consequences.
And do not allow hope, fear, sunk cost or momentum to quietly redefine what “enough” was supposed to mean.
A credible decision does not require every uncertainty to disappear.
It requires the remaining uncertainty to be visible, proportionate and consciously accepted.
Sometimes the evidence will justify proceeding.
Sometimes it will reveal a stronger version of the idea.
Sometimes it will show that the organisation needs to pause.
And sometimes it will provide the clarity required to stop.
Each of those outcomes can represent progress.
Because the purpose of evidence is not to force an idea forward.
It is to help determine what the idea deserves next.
Continue Exploring
This article is Part Seven of Before You Commit, a Cultural Intelligence Studio series exploring evidence, assumptions, constructive challenge, cultural intelligence, scenario thinking, experimentation and better decision-making under uncertainty.
Previous Article
Imagine It Failed
Why a Pre-Mortem Can Reveal What Optimism Leaves Hidden
Article Six explored the discipline of imagining that an idea has already failed, then working backwards to identify hidden vulnerabilities, dependencies, cultural risks and warning signs before they become expensive.
Next Article
Keep the Door Open
Why Reversible Decisions Create More Room for Learning
Article Eight examines how staged commitments, options, boundaries and exit conditions can reduce the cost of uncertainty. It asks which decisions can be delayed, divided, tested or reversed—and why the strongest next step is sometimes the one that preserves the greatest capacity to learn.
Explore the Cultural Intelligence Simulation Lab
The CIS Idea Simulation Review helps founders, artists, creative entrepreneurs, cultural and community organisations, and small businesses examine an idea before substantial commitment.
It brings together:
the Evidence Ledger;
load-bearing assumptions;
cultural lenses;
Red Team challenge;
scenario analysis;
pre-mortem thinking;
the smallest useful test;
decision thresholds;
and a personalised Decision Report leading towards Proceed, Revise, Pause or Stop.
Cultural Intelligence Studio
Human judgement. Cultural intelligence. Strategic clarity. AI-enabled capability.