Skip to article content

Confidence is not competence: When employees are certain but wrong

Confidence can add useful diagnostic context, but it is not proof of capability. Learn how to combine certainty with evidence and improve calibration.

By Alireza Ibrahimi

9 min read

Research Brief
Learn how confidence and accuracy can diverge, what high-confidence errors reveal, and how to build safer, evidence-based workplace assessments.

Confidence is a judgment.

Competence is a demonstrated capability.

The two can align, but they are not interchangeable. An employee may be confident because a procedure is familiar, because a previous employer used a different rule, or because the task appears simpler than it is. Another employee may perform accurately while remaining uncertain because the process is new.

Confidence therefore becomes useful when it is compared with evidence—not when it replaces evidence.

Key takeaways

  • Confidence should refer to a specific answer, decision, or performance—not a person's general worth or intelligence.
  • Combine confidence with accuracy, reasoning, task evidence, and source version.
  • High-confidence errors can reveal a stable competing belief and create a strong correction opportunity.
  • A single confidence rating should not determine certification, routing, or remediation.
  • Calibration can improve when learners repeatedly compare predictions with actual performance and clear standards.
  • Avoid popular labels such as “Dunning-Kruger” when the evidence only shows one confident wrong response.

Confidence answers a different question

A quiz score asks:

Was the response correct according to the scoring standard?

A confidence rating asks:

How certain was the learner that the response was correct?

A demonstration asks:

Could the learner perform the task under the observed conditions?

A workplace measure asks:

Did the behavior or performance occur in the relevant environment?

These are different forms of evidence.

The term calibration is often used for the relationship between confidence and accuracy. Good calibration means that a person's confidence judgments appropriately reflect the likelihood that their responses or performances are correct. A review by Stone describes calibration as the accuracy with which people assess confidence in their own knowledge and connects it to self-regulated learning. Source: Stone, “Exploring the Relationship Between Calibration and Self-Regulated Learning”.

Calibration is not the same as always having low confidence. The goal is not doubt. It is appropriately placed confidence.

Combine confidence with accuracy

A simple two-signal view creates four useful patterns.

Correct and confident

The learner may possess stable knowledge.

Confirm with a different case or task when the consequence of error is high. Familiarity with one item is not the same as flexible capability.

Correct and uncertain

The learner may understand the material but lack fluency, may have guessed, or may not trust the judgment.

Ask for reasoning and provide another opportunity. Targeted feedback can strengthen appropriate confidence without repeating the entire course.

Incorrect and uncertain

The learner recognizes uncertainty.

Provide the missing explanation, a worked example, or guided practice. Also make reliable support easy to access at work.

Incorrect and confident

The learner may be using a stable wrong rule, an outdated procedure, or a misread cue.

Investigate the reasoning and current source. This pattern deserves attention because the learner may not independently seek help.

These patterns are hypotheses for instructional action, not diagnoses. One item can be ambiguous, guessed, or mis-scored.

Four visual states represent correct and confident, correct and uncertain, incorrect and uncertain, and incorrect and confident responses.

Ask for confidence at the right level

Avoid broad prompts such as:

How confident are you in incident reporting?

A person may be confident in routine cases and uncertain about rare exceptions. General self-ratings mix many capabilities and contexts.

Prefer item-level or task-level prompts:

How certain are you that this event requires immediate escalation?

How confident are you that this report meets all required documentation criteria?

How certain are you about the first action in this scenario?

Use a simple scale that employees can understand consistently. The exact number of points matters less than clear anchors.

For example:

  • not sure;
  • somewhat sure;
  • very sure.

Or:

  • I would verify before acting;
  • I believe this is correct but would like confirmation;
  • I am prepared to act and explain the reason.

The second version connects confidence to behavior, but it should still be tested for clarity.

Reasoning makes confidence more interpretable

A confidence rating without reasoning can tell you that a response deserves a closer look. It cannot tell you why.

After an important decision, ask:

  • Which rule applies?
  • What detail drove the choice?
  • What alternative did you reject?
  • What condition would change the answer?
  • Where would you verify this?
  • Which source version are you using?

Consider two employees who both choose the wrong answer with high confidence.

Employee A explains an outdated version of the SOP accurately.

Employee B says the answer “just feels right” and cannot identify a rule.

The same confidence-and-accuracy pattern points to different responses:

  • Employee A needs explicit change-focused retraining and removal of outdated cues.
  • Employee B may need foundational concepts, modeled reasoning, and guided practice.

High-confidence errors can be unusually teachable

Research on the hypercorrection effect has often found that high-confidence errors are more likely than low-confidence errors to be corrected after feedback.

One proposed explanation is surprise: feedback that violates a strong expectation may attract additional attention. Fazio and Marsh found that surprising feedback was better remembered in experimental general-knowledge tasks. Source: “Surprising Feedback Improves Later Memory”.

But correction can be temporary.

Butler, Fazio, and Marsh found that the hypercorrection effect persisted after a one-week delay, while overall correction declined. When the correction was forgotten, the original high-confidence error was more likely to return. Source: “The Hypercorrection Effect Persists Over a Week”.

The practical implication is not that a confident error fixes itself.

It is:

Use the moment of surprise for clear explanatory feedback, then require retrieval and application again after time has passed.

These experiments used general-knowledge questions, not high-risk workplace procedures. Local validation and stronger performance evidence remain necessary.

Do not use confidence to shame or rank people

A confident error can be funny in a meme and dangerous in a learning culture.

If employees expect ridicule, public exposure, or punishment, they may:

  • choose artificially low confidence;
  • avoid answering;
  • search for the answer instead of revealing current knowledge;
  • hide uncertainty at work;
  • become defensive rather than examining the reasoning.

Keep confidence checks low stakes during learning. Explain why they are being collected and how the information will be used.

For formal certification or employment decisions, use validated instruments, job-relevant performance standards, appropriate governance, and qualified review. A casual confidence question is not a defensible selection test.

Avoid the Dunning-Kruger shortcut

The phrase “Dunning-Kruger effect” is often used online to describe any confident person who is wrong.

That is an overreach.

A single high-confidence error does not establish a general metacognitive pattern, and it says nothing about the person's intelligence or competence outside the specific task. Performance, self-estimation, task difficulty, measurement design, and statistical artifacts all complicate broad conclusions.

For workplace learning, the more useful language is specific:

The learner expressed high confidence in an incorrect response to this scenario.

Then gather more evidence.

Precise description supports action. A popular psychological label often creates judgment without improving the intervention.

Build calibration through standards and feedback

People can compare confidence with performance only when performance standards are visible.

Provide:

  • clear criteria;
  • examples of acceptable and unacceptable work;
  • explanation of why each example meets or misses the standard;
  • opportunities to predict performance;
  • feedback comparing prediction with evidence;
  • new tasks to see whether judgment improves.

An experimental study by Nederhand and colleagues examined whether performance standards could improve students' calibration on current and similar tasks. The findings support the broader idea that standards and feedback can help people judge performance more accurately, while effects may differ by performance level. Source: Nederhand and colleagues, “Learning to Calibrate”.

A later classroom study found that students' confidence calibration broadly improved across a sequence of six low-stakes mathematics assessments. Source: Foster and Renie, “Changes in Students’ Confidence Calibration”.

These are educational contexts, not direct evidence that the same routine will improve workplace decision calibration. A workplace team can nevertheless test a similar loop:

predict → perform → compare with standard → receive feedback → try a new task

Use confidence differently by risk

The consequences of miscalibration vary.

For a low-risk administrative task, an incorrect-confident response may trigger a brief explanation and another scenario.

For a safety-critical or regulated task, it may trigger:

  • immediate targeted remediation;
  • supervised practice;
  • confirmation of the current source;
  • delayed reassessment;
  • manager observation;
  • temporary requirement to verify before independent action.

Do not create false precision from a confidence number. “90% confident” does not mean a 90% probability of safe performance unless the scale and interpretation have been properly studied.

Use confidence as one signal in a larger evidence system.

Measure capability beyond the quiz

Confidence attached to multiple-choice answers cannot establish competence in a procedure.

Depending on the outcome, combine it with:

  • explanations;
  • realistic scenarios;
  • task demonstrations;
  • work samples;
  • simulations;
  • delayed retrieval;
  • workplace observations;
  • operational quality measures.

A capable employee should also know when confidence is insufficient and verification is required.

That is an important outcome:

The employee recognizes uncertainty, accesses the approved source, and escalates when the situation exceeds their authority.

Competence sometimes means acting decisively. It sometimes means knowing not to act alone.

How this can shape LoreGraph

LoreGraph helps organizations transform workplace knowledge into learning, practice, and assessment. Confidence could add useful context when it is represented carefully.

A future evidence model might connect:

learner response → confidence judgment → reasoning → source-supported answer → feedback → later response → performance evidence

The system should preserve:

  • the exact question or task;
  • the current source version;
  • whether the response was correct;
  • how confidence was collected;
  • the reasoning supplied;
  • uncertainty in any inferred learner state;
  • the instructional action taken.

Learn more about LoreGraph.

AI should not conclude that a person is “overconfident” from one interaction. It can help identify evidence patterns for human review and generate targeted practice grounded in approved sources.

Limits and common mistakes

Confidence research includes laboratory tasks, school assessments, and other contexts that differ from workplace performance. Results should be applied cautiously.

Avoid these mistakes:

  • Equating confidence with competence: Use direct capability evidence.
  • Treating low confidence as failure: Correct performance may need reinforcement, not remediation.
  • Treating high confidence as arrogance: It may reflect familiarity, an old rule, or unclear standards.
  • Using one item to label a learner: Look for patterns.
  • Collecting confidence with no purpose: Define the action it will inform.
  • Creating false numerical precision: A rating scale is not automatically calibrated probability.
  • Ignoring the source version: A learner may confidently remember what was previously correct.
  • Skipping delayed evidence: The correction may not persist.
  • Using confidence data for high-stakes decisions without validation: Formative signals and employment assessments are not interchangeable.

A learner predicts confidence, performs a task, receives feedback, compares judgment with evidence, and tries a new task.

Next step

Add a confidence prompt to three important scenario questions—not an entire course.

For each response, capture:

  1. the decision;
  2. confidence;
  3. a brief explanation;
  4. the source-supported answer;
  5. targeted feedback;
  6. performance on a new case later.

Review the four patterns and decide which evidence actually changed your instructional response.

The goal is not to make employees less confident.

It is to help confidence become better aligned with current, demonstrated capability.

Sources and further reading


Alireza Ibrahimi

Founder, LoreGraph

Software engineer and Learning Engineering researcher building AI systems that transform workplace knowledge into measurable learning experiences.

Put it into practice

Create structured training with measurable progress

Give your team training with practice, assessment, and progress you can actually see.

Explore team training

Keep reading