The psychology of idea evaluation and why the same minds that generate good ideas are poorly equipped to judge them
The most important finding in creativity research that commercial practice most consistently ignores is not about idea generation. It is about idea evaluation: the cognitive systems that generate genuinely creative ideas are distinct from, and systematically biased against, the cognitive systems that evaluate them. The result is that the most valuable ideas are reliably filtered out at the evaluation stage — not by bad evaluators but by the mechanics of how evaluation works.
The Mueller finding: novelty as a threat signal to the evaluation system
Mueller, Melwani and Goncalo’s (2012) Psychological Science research provided the most direct experimental evidence for the evaluation bias. Participants who were primed to feel uncertain showed more negative implicit associations with creative ideas — even while explicitly stating that they valued creativity. The gap between explicit endorsement and implicit rejection is the mechanism that makes the bias so commercially consequential: the evaluator genuinely believes they are assessing ideas on their merits while a distinct, automatic cognitive process is filtering out the most novel options before deliberate evaluation begins.
The mechanism is System 1’s threat-detection response applied to novelty. Genuinely novel ideas — the ones that are furthest from existing schemas — violate the patterns that the prediction system uses to evaluate outcomes. Schema violation activates the uncertainty-aversion response, which registers as a negative signal before the deliberate evaluation process begins. The evaluator experiences this as a feeling that something is wrong with the idea, or that the idea is risky, or that implementation would be difficult — all of which are plausible post-hoc rationalisations of the prior automatic rejection.
The practical consequence is systematic. The most genuinely creative ideas — the ones that would create most value precisely because they are furthest from existing schemas — will reliably feel most uncertain, most threatening, and most risky to evaluate. The evaluation bias is not against bad ideas; it is against novel ones. And novelty is the property that makes a creative idea valuable.
The expert evaluator problem: Einstellung at evaluation
The Bilalić et al. (2008; 2010) Einstellung research predicts the specific way that expert evaluation amplifies the novelty bias. The expert has deeply developed schemas for what successful solutions in their domain look like — schemas built from experience with what has worked before. These schemas activate automatically during evaluation, directing attention toward familiar-solution territory and away from the genuinely novel alternative.
The more expert the evaluator, the stronger the Einstellung activation. The experienced product manager evaluating a novel product concept is comparing it, automatically and largely unconsciously, to the mental model of successful products they have internalised through years of domain experience. The genuinely novel concept — the one that would succeed through mechanisms the expert schema does not represent — is evaluated against the wrong standard and found wanting.
This explains the well-documented pattern of expert rejection of eventually successful innovations. Airbnb, Uber, and Dropbox were rejected by investors whose expert schemas about what successful marketplace businesses, transportation services, and cloud storage solutions look like did not accommodate the genuinely novel mechanisms these products employed. The expert rejection was not incompetent; it was the Einstellung mechanism operating precisely as the research predicts it will.
The dual process mismatch: System 2 evaluating System 1 outputs
The dual process framework identifies a third mechanism that explains why the evaluation system is poorly calibrated for creative output specifically. Creative ideas are generated by System 1 — the DMN’s automatic, associative, non-directed processing that connects stored knowledge in remote and novel ways. They are evaluated by System 2 — the ECN’s deliberate, analytical, structured processing that applies criteria of logical consistency, feasibility, and risk.
The problem is not with System 2 as an evaluation tool; it is with applying System 2 evaluation criteria to System 1 outputs that have not yet been developed enough to satisfy those criteria. A genuinely novel idea at the moment of generation is not logically consistent — it is a raw associative connection that requires development before its internal logic can be articulated. System 2 evaluating the undeveloped idea against criteria of logical consistency will find it inconsistent and reject it. But the inconsistency is a feature of the developmental stage of the idea, not a feature of the idea’s ultimate quality.
The idea that would become the product needs to be developed sufficiently for the System 2 evaluation to be calibrated appropriately. Evaluating it before it is developed produces the false negative that the Einstellung and Mueller mechanisms predict — the genuinely valuable idea rejected because it cannot yet defend itself in the analytical language that System 2 evaluation requires.
The group evaluation amplifier
The social dynamics of group evaluation add a further layer of distortion to the individual evaluation biases. Asch’s conformity research and Janis’s groupthink research together predict that group evaluation of novel ideas will be more conservative than individual evaluation — because the uncertainty-aversion response that novel ideas trigger in one group member signals uncertainty to other group members, activating the normative social influence that produces conformity toward the negative evaluation.
The group evaluation setting therefore amplifies the individual evaluation bias: the first negative response signals uncertainty to the group; the uncertainty signal activates the same uncertainty-aversion in other members; the convergent negative evaluation is reinforced across the group through the social proof mechanism; and the novel idea is rejected with a consensus that feels like considered judgment but reflects the social dynamics of shared uncertainty-aversion.
Amabile’s consensual assessment technique research established that genuine creative quality can be accurately evaluated by domain experts — but only when the evaluation is independent, separated from social dynamics, and protected from the conformity mechanisms that group evaluation activates. The technique requires individual evaluation before group discussion, not group evaluation replacing individual assessment.
What the research supports as better evaluation practice
The research implies specific structural interventions for idea evaluation. Separating generation from evaluation — never evaluating ideas in the same session in which they are generated — protects the generative phase from the premature convergence that evaluation activates and provides the developmental period that novel ideas require before System 2 evaluation is calibrated appropriately.
Evaluating ideas individually before group discussion — the Amabile consensual assessment technique applied to commercial idea evaluation — prevents the social conformity dynamics from contaminating independent expert judgment. The individual assessments then inform the group discussion rather than the group dynamics determining the individual assessments.
Assessing for potential as well as feasibility — explicitly separating the question “what would this idea need to become viable?” from the question “is this idea currently viable?” — preserves the options that the Mueller uncertainty-aversion would otherwise eliminate at first encounter. The genuinely novel idea that fails the feasibility test may pass the potential test; combining the two questions into a single evaluation produces the false negative that feasibility-focused expert evaluation reliably generates.
Books worth reading on this
The Runaway Species by David Eagleman and Anthony Brandt is the most accessible available neuroscience account of how human creative evaluation actually works — covering the specific brain mechanisms through which novelty is assessed, why the evaluation system has a systematic conservative bias, and what the most creatively productive human enterprises have done to counteract it. For the entrepreneur who wants the most readable available treatment of why the evaluation system works against the most valuable ideas, Eagleman and Brandt provide the most engaging available account.
If the dynamics described here are significantly affecting your wellbeing, speaking with a psychologist is the right next step. UK: Samaritans (116 123, free, 24/7). Mind (0300 123 3393). BACP: bacp.co.uk/search/Therapists. Crisis Text Line — text HOME to 741741 (US, UK, Canada, Ireland). International: internationaltherapistdirectory.com.
This article is for educational and informational purposes only. Sources: Mueller, J.S., Melwani, S. & Goncalo, J.A. (2012), The Bias Against Creativity: Why People Desire but Reject Creative Ideas, Psychological Science, 23(1), 13–17. Bilalić, M., McLeod, P. & Gobet, F. (2008), Why Good Thoughts Block Better Ones, Cognition, 108(3), 652–661. Bilalić, M., McLeod, P. & Gobet, F. (2010), The Mechanism of the Einstellung (Set) Effect, Current Directions in Psychological Science, 19(2), 111–115. Kahneman, D. (2011), Thinking, Fast and Slow, Farrar, Straus and Giroux. Amabile, T.M. (1983), The Social Psychology of Creativity, Springer. Eagleman, D. & Brandt, A. (2017), The Runaway Species, Catapult. Kahneman, D., Sibony, O. & Sunstein, C.R. (2021), Noise: A Flaw in Human Judgment, Little, Brown.
Have a Question?
Submit your question and we may cover it in a future article.