The Mean-Variance Innovation Tradeoff in Evaluations
Our two field experiments with 2,723 evaluators revealed that screening for feasibility first produces a stronger set of selections on average, while screening for novelty first protects the rare breakthroughs.
By Cyrille Grumbach, Jackie Lane, and Georg von Krogh
July 2026

Key takeaway: The criterion you apply first determines which solutions remain in contention at all. A criterion applied second can only reject; it cannot recover what the first screen already discarded. Lead with feasibility for a reliably strong portfolio, lead with novelty to keep moonshots alive, and use AI assistance when you want to raise the average without letting weak ideas through.
This project looks at a stage of innovation that often gets less attention than idea generation: evaluation. When an organization or a contest receives hundreds or thousands of submissions, the choices made during screening shape both the rate and the direction of innovation. At the center of that process sits a familiar tension between novelty and feasibility. Evaluators cannot give equal weight to both criteria at the same moment, so in practice they sequence them, prioritizing one criterion at a first stage and the other at a second. What has not been clear is whether that sequence is a neutral procedural detail or a decision that quietly determines the outcome.
To find out, the project ran two field experiments with 2,723 evaluators, producing 54,460 solution–evaluator pairs. Evaluators were randomly assigned to one of two sequences. Some screened first for feasibility and then for novelty. Others screened first for novelty and then for feasibility. Each sequence was also run with and without AI assistance, so the project could test whether an AI partner changes the dynamics of the tradeoff.
The results show what the project calls a mean–variance innovation tradeoff. A feasibility-then-novelty sequence produced selections with higher average innovation. A novelty-then-feasibility sequence produced greater variance in innovation. Neither sequence is simply better. One delivers a dependable portfolio, the other delivers a wider spread that includes both weaker picks and more extreme upside.
AI assistance did not eliminate this pattern. The tradeoff persisted whether or not evaluators had AI support. What AI assistance did do was shift the whole distribution: it raised the mean and compressed the variance, and it helped evaluators catch more "stars," the solutions that are both highly novel and highly feasible. In other words, AI improved the quality of screening without resolving the underlying tension between the two criteria.
The project also shows when sequencing matters and when it does not. When the two criteria are aligned, meaning a solution scores high on both or low on both, the order made little difference. Easy cases stayed easy. Sequencing became decisive when the criteria diverged. A novelty-then-feasibility sequence preserved more "moonshots," the atypical solutions that are only barely feasible. A feasibility-then-novelty sequence favored "safe bets," solutions that are workable but conventional. The order, in effect, chooses which kind of imperfect solution gets a second look.
The evaluators' own stated reasons explain the mechanism. At the first screen, they invoked their assigned criterion to justify both the solutions they accepted and the ones they rejected. At the second screen, the assigned criterion shaped only their rejections. A criterion applied second therefore operates through exclusion rather than selection. It can remove candidates from the surviving pool, but it cannot reach back and rescue anything the first screen already eliminated. Whichever criterion leads determines the set of solutions that remain in play.
For practice, the implication is concrete. Sequencing is not an administrative detail to be settled by convenience or habit; it is a strategic lever that should follow from the goal of the evaluation. If the objective is a solid slate of implementable solutions, put feasibility first. If the objective is to find the rare idea that changes the field, put novelty first and accept the wider spread that comes with it. And if the aim is to lift overall quality, AI assistance helps, but it will not choose the direction for you. That choice is still made the moment you decide which criterion goes first.

© 2026 Build2Gether. All rights reserved.