🎧 Listen to this article

An AI-generated podcast companion to this article, produced with Google NotebookLM.

The traits that get someone chosen as a leader and the traits that make someone effective once they're in the role are only loosely related. Most organizations have never noticed the gap, because the selection process is built to reward the first and structurally blind to the second until the damage is already done.

Here is a pattern that has played out, in some version, at nearly every organization large enough to have a hiring committee or a board. A candidate arrives with an unmistakable presence. They are decisive in the room, fluent under pressure, quick with a compelling narrative about where the business needs to go. Everyone who meets them comes away with the same read: this is a leader. They are hired, often at a senior level, sometimes the most senior level available. Eighteen months later, or three years, or five, the wreckage surfaces. Departed talent. A team nobody wants to join. A pattern of credit taken and blame assigned that finally becomes too visible for the organization to keep looking past.

The diagnosis people reach for afterward is almost always personal: we hired a bad apple, or the culture protected someone it shouldn't have. That explanation is comfortable because it treats the failure as an exception. It rarely is. It is usually the selection system working exactly as designed, just not for the outcome anyone intended, because the system was built to detect confidence and charisma, and confidence and charisma are a measurably different thing from competence.

The Cost Nobody Line-Items

Elizabeth Holmes did not talk her way past Theranos's board with technical credentials. She talked her way past it with presence, a black turtleneck, a voice reportedly pitched well below her natural register, a story about revolutionizing blood testing so compelling that a board that included a former secretary of state and a former secretary of defense signed on without demanding the technical verification a much lower-stakes hire would have required. The company raised hundreds of millions of dollars before a 2015 Wall Street Journal investigation revealed that its core technology could not do most of what it claimed. Holmes was convicted of investor fraud in 2022. The story gets told as an outlier. It is better understood as a magnified version of an ordinary process: a group of experienced, intelligent people mistook a fluent performance of competence for the thing itself, and none of the normal checks caught it in time.

Adam Neumann is the same pattern with the correction happening faster. WeWork was valued at 47 billion dollars in early 2019, on the strength of a founder whose pitch was closer to a spiritual movement than a real estate business, and on a governance structure that gave him outsized voting control while the company leased buildings he personally owned. Almost none of that was hidden. It was disclosed, in writing, in the S-1 filing WeWork submitted ahead of its planned IPO. The moment outside investors were asked to read the filing rather than sit in the room with Neumann, the valuation collapsed by more than three-quarters within months, and Neumann was pushed out by his own board. The charisma had not changed. What changed was the format of the evaluation, from an impression formed in person to a set of facts on a page. That is a small experiment run at enormous scale, and it has a clean result: take away the room, and the story stops working.

Uber's Travis Kalanick shows the version of this that plays out over years instead of months. Kalanick's aggressive, take-no-prisoners style was treated internally as the engine of Uber's growth, and the culture that style produced was tolerated well past the point where it was generating obvious harm, until a former engineer named Susan Fowler published an account of the discrimination and harassment she experienced there. The resulting investigation, led by former U.S. Attorney General Eric Holder, found conduct serious enough to produce more than twenty firings and, eventually, Kalanick's own resignation under pressure from Uber's largest investors. None of what the investigation found was new information to the people who worked there. It had simply never been asked for, because the leadership style producing it was also the style being rewarded.

The research explains why this keeps happening rather than being a string of bad luck. A 2010 study of over two hundred corporate professionals in management-development programs, led by the psychologist Paul Babiak alongside Craig Neumann and Robert Hare, found that psychopathy in executives correlated positively with in-house ratings of charisma, creativity, and communication skill, and correlated negatively with ratings of actual performance, teamwork, and management ability. Style and substance moved in opposite directions in the same people, rated by the same organizations. A separate 2015 meta-analysis of the research on narcissism and leadership, led by Emily Grijalva, found that narcissism reliably predicts who gets chosen as a leader, but shows no relationship, or even a negative one past a certain point, with how effective that person actually turns out to be once independent observers are the ones rating them. And a widely cited study of CEO narcissism by Arijit Chatterjee and Donald Hambrick found that narcissistic CEOs don't simply perform worse, they produce more volatile performance: bigger swings, bigger bets, more extreme outcomes in both directions, which is a much harder pattern for a board to catch early than simple underperformance would be.

Why the System Keeps Getting Fooled

None of this is because hiring committees are careless. It's because the moment that matters most in most selection processes, the interview, the pitch, the board presentation, is precisely the moment where charisma and competence are hardest to tell apart. Confidence, fluency, and decisiveness are real signals of real leadership ability. Most of the personality research on effective leaders finds a moderate, genuine relationship between assertiveness and leadership effectiveness. But that same surface behavior is also exactly what a narcissistic or exploitative operator is best at producing on demand, for the length of a single high-stakes performance. An interview panel is not measuring the ability to lead. It's measuring the ability to be persuasive for sixty minutes, and assuming the two are the same thing.

This gets compounded by a well-documented feature of how organizations judge leadership more broadly. A 2011 meta-analysis by Amy Koenig, Alice Eagly, and colleagues, pulling together decades of research across three separate research traditions, found that the implicit template most people use to recognize "a leader" is still substantially masculine and take-charge, which means the presence being rewarded in that sixty-minute window isn't even a neutral read of competence, it's a read against a specific stylistic prototype. That has two consequences worth separating. It means organizations are vulnerable to being fooled by anyone fluent in that particular style, whether or not they have the substance behind it. And it means people who lead differently, more quietly, more collaboratively, without a taste for self-promotion, get underweighted by the same evaluation twice: once for not resembling the prototype, and again because nobody is checking whether the prototype was ever a good proxy for performance in the first place.

How to Change the Rules, If You're the One Doing the Choosing

The organizations that get this right don't rely on better instincts in the interview room. They change what the room is measuring.

Google's Project Oxygen is the clearest example of what that looks like done well. In 2009, Google's people-analytics team set out to test an assumption almost everyone in the company held, that the best managers were the most technically brilliant ones. They analyzed performance reviews, exit interviews, and survey data across thousands of managers to find out what actually separated the best from the rest. Technical expertise came in last of the eight qualities identified. What came first was almost entirely relational: coaching people well, listening, showing genuine interest in a report's growth and well-being, being available rather than commanding. Google didn't discover this by trusting its instincts about what good leadership looks like. It discovered it by refusing to trust those instincts, and building a measurement system that could contradict them. That is the model. Any organization willing to run the equivalent exercise, testing its own promotion criteria against its own outcome data, will very likely find the same gap between the traits it has been rewarding and the traits that actually predict performance.

Underneath that, a few concrete changes carry real evidence behind them. Structured interviews, ones built around a fixed set of job-relevant questions and a defined scoring rubric rather than open conversation, are meaningfully more valid predictors of performance than unstructured ones, and they close much of the room where pure charisma gets to operate unchecked, according to a comprehensive 2014 research review by Julia Levashina and colleagues. Reference checks are worth redesigning with the same intent: rather than asking a candidate's chosen references whether they'd rehire the person, which almost everyone answers generously, ask for a specific instance of how the candidate treated someone with less power or status than them, an assistant, a junior hire, a vendor. That single question surfaces the entitlement and exploitation patterns that charisma is specifically good at hiding, and it is far harder to spin than a character reference. For internal promotions, pull multi-rater feedback before the decision, not after, and pay particular attention to the gap between how someone rates their own performance and how their peers and direct reports rate it. That gap, more than either number alone, has real predictive value, and a wide one is one of the more reliable early warning signs available. And formalize sponsorship, deliberately connecting rising talent with a senior advocate who will vouch for them before they're in the room, so that advancement isn't gated entirely by how someone performs in a single high-stakes impression.

What to Do If You Don't Have the Natural Style

None of this changes fast, and if you're the one being evaluated rather than the one doing the evaluating, waiting for the system to correct itself is not a strategy. Five things are worth doing in the meantime, and none of them require becoming someone you're not. They happen to spell TABLE, which is a fitting word for what this whole piece has been about: whose table it is, who gets to define what belongs at it, and how you earn a seat without losing yourself in the process.

T
Track Record Change what you're being judged on — build proof that's harder to wave away than charisma
A
Assert With Warmth Pair genuine assertiveness with visible warmth to reduce the backlash quieter leaders face
B
Build Substance, Not Stagecraft Invest in decisiveness and follow-through, not vocal register or airtime
L
Line Up Sponsorship Seek advocacy, not just mentorship — someone willing to vouch for you before you're in the room
E
Evaluate the Table Be selective about where you spend your energy — not every table has the same rules

T — Track Record

Change what you're being judged on whenever you have any say in it. Impression-based evaluation is where style-over-substance bias does the most damage. A track record built on numbers other people can verify, a result, a project, a number that moved, is much harder to wave away with charisma, and it's available to introverted people by design, since it doesn't require winning the room to be true.

A — Assert With Warmth

Pairing genuine assertiveness with visible warmth, rather than either alone, measurably reduces the social penalty that quieter or non-traditional leaders, and take-charge women in particular, tend to face for stepping into the room fully. Research on workplace gender bias by Laurie Rudman and Peter Glick found exactly this pattern: assertiveness alone invites backlash, assertiveness paired with warmth mostly doesn't.

This isn't theoretical. When Alan Mulally took over a Ford Motor Company projected to lose seventeen billion dollars, he ran famously disciplined weekly meetings where every executive had to report project status in plain red, yellow, or green terms. For months, everyone reported green while the company hemorrhaged money, until one executive finally marked a project red. Mulally didn't reprimand him. He applauded and asked the room who could help. That single moment, hard accountability paired with warmth toward the person who told him the truth, is widely credited with turning Ford's culture from one of fear into one of candor. Indra Nooyi, PepsiCo's CEO for twelve years, built a quieter version of the same habit: she wrote hundreds of personal letters a year to the parents of her senior executives, thanking them for their child's contribution. Neither leader was any less demanding for it.

Both point to the same practical takeaway, and it doesn't require a grand gesture: naming a specific contribution out loud instead of a generic thank-you, remembering a detail about someone's life and following up on it later, and meeting someone's honesty about a mistake with curiosity instead of punishment. That's not code-switching, the practice of shifting how you talk, present, or carry yourself to fit a room that wasn't built with people like you in mind. Research by Courtney McCluney and colleagues on the psychological costs of code-switching has found it carries a real cognitive tax, a steady drain on the mental energy that would otherwise go toward the work itself, along with a documented link to burnout. It's leading with the whole of a real repertoire instead of half of it, not suppressing part of yourself to get in the room.

B — Build Substance, Not Stagecraft

Be precise, in your own development, about what's actually worth building. Confidence as theater, a certain vocal register, forced enthusiasm, dominating the airtime in a meeting, has close to no relationship to real leadership effectiveness. Decisiveness, directness, and a willingness to make the call and own it, do. Spend your effort on the second category. It's a real, learnable, well-validated skill, and unlike the theater, it doesn't require you to perform a version of yourself you can't sustain.

L — Line Up Sponsorship

Actively seek it rather than waiting to be noticed. Research by Herminia Ibarra and colleagues on why women are promoted less often than equally qualified men found they are frequently over-mentored and under-sponsored, given advice but not advocacy. A senior person willing to vouch for you before you're being evaluated changes the terms of the evaluation itself, and it is one of the highest-leverage moves available to anyone who would rather be measured on results than on stage presence.

E — Evaluate the Table

Be selective about where you spend your energy fighting this battle at all. The masculine, take-charge leadership prototype isn't equally rigid everywhere; the same research on leadership stereotypes found it's noticeably less pronounced in some organizations and fields than others. Not every table has the same rules, and it is entirely reasonable to choose the ones where the rules already look more like the ones you'd want to be judged by.

The Question Worth Asking Before You Promote Anyone

Every example in this piece shares the same root cause: an organization mistook the ability to describe results for the ability to produce them, because describing results well is fast, visible, and happens in the room, while producing them is slow, often invisible, and happens somewhere else entirely.

The question worth asking before any hiring or promotion decision is not "did this person impress us." It's "what is the evidence, independent of how this person describes themselves, that they can do what the role actually requires." That single substitution, evidence for impression, is a harder discipline to build than it sounds, because impression is free and immediate and evidence takes real effort to gather.

The Question

What is the evidence, independent of how this person describes themselves, that they can do what the role actually requires?

But it is the only version of the question that has ever reliably told the difference between a leader and a good performance of one.