Data & ROI

Measuring managerial skills objectively

Every manager has been evaluated at least once by a superior, a peer, or a consultant. And every time, the same problem arises: the evaluation says as...

Fionnuala O'Driscoll
8 min read
évaluationcompétences managérialesassessment

1. The bias problem with traditional evaluations

Every manager has been evaluated at least once by a superior, a peer, or a consultant. And every time, the same problem arises: the evaluation says as much about the evaluator as the person being evaluated.

Here are the main biases documented by organizational psychology research:

The halo effect. Marie gives excellent presentations to the executive committee. Her manager unconsciously concludes that she must also be excellent at conflict management, active listening, and delegation. One visible competency contaminates the assessment of all the others.

Recency bias. Thomas had a difficult quarter in January-February but excelled in March. His six-month review will only reflect the last two weeks. Six months of data reduced to a recent memory.

Leniency or severity bias. Some managers rate everyone 4 out of 5 — out of comfort, fear of conflict, or culture. Others systematically rate low, convinced that "nobody's perfect." The score reflects the evaluator's philosophy, not the employee's reality.

Similarity bias. We rate people who resemble us more favorably — same communication style, same background, same interests. It's documented, reproducible, and nearly impossible to consciously eliminate.

Attribution bias. When an employee I like fails, it's because "the situation was difficult." When an employee I have a tense relationship with fails, it's "a competence issue." Same situation, two opposite interpretations.

The result? Traditional evaluations measure the quality of the relationship between evaluator and evaluated more than actual competence. They measure impressions, not behaviors.


2. What gets measured: 5 behavioral dimensions

The system doesn't give an "overall impression." It analyzes observable, quantifiable behaviors within the precise context of a simulated conversation. Here are the five main dimensions.

Active listening

What it is: The ability to demonstrate that you've heard and understood before responding.

What the system detects: Does the person rephrase what the other person said before giving their own response? Do they ask open-ended questions to dig deeper? Do they acknowledge expressed emotions?

Metrics: Rephrasing rate, open-ended to closed-ended question ratio, emotional acknowledgment frequency.

Assertiveness

What it is: The ability to express a clear, direct position without tipping into aggression.

What the system detects: Are messages clear and structured? Does the person state expectations without ambiguity? Do they use direct formulations without aggression markers?

Metrics: Clarity score, directness index, presence of aggression markers (ultimatums, threats, sarcasm).

Emotion management

What it is: The ability to recognize emotions (one's own and others') and regulate tension.

What the system detects: Does the person name the emotions they perceive? Do they escalate or de-escalate tension? Do they stay regulated when the simulated counterpart increases pressure?

Metrics: Emotional recognition rate, escalation/de-escalation patterns, tone stability under pressure.

Solution orientation

What it is: The ability to generate options and build collaborative solutions rather than imposing or blocking.

What the system detects: How many solution pathways does the person propose? Do they seek to build with the other person or impose their vision? Do they facilitate co-construction?

Metrics: Number of options generated, collaborative vs. directive ratio, presence of co-construction language.

Communication style

What it is: The register adopted and its consistency with the situational context.

What the system detects: What register does the person use (authoritative, empathetic, neutral, coaching)? How does it evolve throughout the conversation? Do they adapt to context?

Metrics: Register classification, lexical diversity, adaptation patterns throughout the exchange.


3. How it works (simplified version)

You don't need to be a data scientist to understand the mechanism. Here's the process in four steps.

Step 1: Natural Language Processing (NLP). Every message sent during the simulation is analyzed by a natural language processing model. The system identifies sentence structure, vocabulary used, and the intentions behind each intervention.

Step 2: Behavioral marker detection. The system looks for specific patterns: a rephrasing, an open-ended question, an emotional acknowledgment, an aggression marker, a solution proposal. Each marker is associated with one of the five dimensions.

Step 3: Sentiment and tone analysis. Sentiment analysis detects emotional shifts throughout the conversation. Is the tone escalating? Is the person staying stable? Was there a successful de-escalation after a tension peak?

Step 4: Longitudinal comparison. This is perhaps the most powerful advantage. The system compares your results with your previous simulations. It detects trends: Is your active listening improving? Does your assertiveness drop in authority scenarios? Machine learning refines accuracy over time.

In summary: Conversation → NLP analysis + behavioral markers → Scores per dimension + personalized recommendations.


4. The real advantages

Consistency

The system applies the same criteria to every evaluation, every time. No good-mood days or bad-mood days. No "I'm more demanding on Friday afternoons." Standards are identical for the first employee evaluated and the thousandth.

Objectivity

No personal relationship biasing the judgment. the system doesn't know if you're likeable, if you share the same background, or if you invited it to lunch last week. It analyzes behaviors, period.

Granularity

Where a human evaluator gives an "overall impression," the system measures more than 12 specific behaviors. You don't receive "you're good at communication" — you receive "your rephrasing rate is 34%, up 8 points from last month, but your emotional recognition remains below the median."

Longitudinal tracking

The system doesn't measure a single point in time — it traces progression curves over weeks and months. You can see concretely where you're improving and where you're plateauing.

No social pressure

Facing a human evaluator, we manage our image. We modulate our speech. We do "impression management." Facing the assessment system, that pressure disappears. The observed behaviors are more authentic.

Scalability

Evaluating 10 managers or 10,000 managers with exactly the same rigor, the same level of detail, the same objectivity. That's impossible for a consulting firm. It's trivial for an automated system.


5. The honest limitations

The non-verbal blind spot

Body language, micro-expressions, tone nuances — all of this escapes text-based analysis. Voice the system is progressing but still doesn't capture posture, gaze, or gestures. In a real conflict, non-verbal communication represents a significant portion of the message.

Cultural context

Communication norms vary considerably across cultures. What is "assertive" in France may be perceived as "aggressive" in Japan or "too indirect" in the United States. the system models require cultural calibration that few solutions integrate today.

Relational intelligence

The quality of a human relationship — trust built over months, shared history, mutual understanding — cannot be measured through an isolated simulation. the system measures behaviors in an exercise, not the depth of a relationship.

Existential depth

The system can measure that your assertiveness drops in authority scenarios. It cannot explore why — your beliefs about power, your fears, your personal history with authority figures. That work remains deeply human.

The gaming risk

People could learn to "check the boxes" — mechanically rephrasing, reflexively asking open-ended questions, without genuine skill development. the system measures observable behaviors, not necessarily the sincerity behind them.


6. The hybrid approach: the system + human coach

The real power lies neither in the system alone nor in human coaching alone. It lies in the combination.

The system identifies patterns. It detects that your assertiveness systematically drops in scenarios involving authority conflict. It spots that your active listening is excellent at the beginning of a conversation but degrades under pressure. It measures objectively.

The coach interprets. They take that data and contextualize it: "Your assertiveness drops when facing authority — let's explore why. What happens emotionally when your counterpart raises their voice?" The coach goes where the system cannot: into understanding meaning.

The system verifies impact. After coaching sessions, the system measures whether behaviors have actually changed. Did the coaching produce observable results? This is the missing link in most development programs.

The best of both worlds: the rigor of data + the wisdom of human guidance.


Key takeaways

"Automated assessment doesn't replace human judgment — it gives human judgment better data to work with."

  • Traditional evaluations reflect the evaluator more than the evaluated (only 20-30% of variance tied to actual performance)
  • the system measures 5 precise behavioral dimensions: active listening, assertiveness, emotion management, solution orientation, communication style
  • Key advantages are consistency, objectivity, granularity, and longitudinal tracking
  • Limitations exist: non-verbal, cultural context, relational and existential depth
  • The optimal approach combines the system measurement with human interpretation

Discover your managerial profile

TrustLeader measures 12 dimensions of your managerial style through realistic conflict simulations. Objective, precise, bias-free.

Create my free account


Related articles

Back to blog