Edition #40: The Job Nobody Is Counting
Inside AI evaluation: the fastest-growing career in tech, why your company already needs one, and what it takes to do the work
By Mykel Salomon ยท 2026-10-02
<hr><h3><strong>๐ Opening Reflection</strong></h3><p>Every week brings another headline about the jobs AI is taking. Far fewer people are asking the obvious follow-up question: who checks the machines?</p><p>Somebody has to. Somebody has to decide what an AI system should and should not do, design the tests that prove it, probe it for the ways it breaks, and sign off before it touches a customer, a patient, or a loan application. That is not a hypothetical role. It is a profession, it is hiring right now, and most people have never heard of it.</p><p>My guest this week is <a target="_blank" rel="noopener noreferrer nofollow" class="text-[#4FD1C5] underline hover:opacity-80 ember-view" href="https://www.linkedin.com/in/malakumar/"><strong>Mala Kumar</strong></a> , Executive Director of <a target="_blank" rel="noopener noreferrer nofollow" class="text-[#4FD1C5] underline hover:opacity-80 foSeUQNNUAlOofCgTsBsMkWwskcxammaBWdAw " href="https://www.linkedin.com/company/humaneintelligence/"><strong>Humane Intelligence</strong></a> , with a career spanning the United Nations, GitHub, the World Health Organization, and public interest technology. Her organization is building one of the most important new categories of this era: human-centered AI evaluation.</p><p>Her reframe is the one I keep returning to. We are no longer only in the era of training humans. We are in the era of <strong>training the machine.</strong> And if you can get your head around that shift, a lot of career paths open up that did not exist three years ago.</p><p>This edition is a practical map of that work. What the job is. Why security breaches have made it urgent. What companies get wrong. What skills actually matter. And what it pays.</p><p>Because AI systems do not need humans in the loop as a formality. They need humans because humans understand consequences.</p><hr><h3><strong>๐ฆ Signal of the Week</strong></h3><p>Three sets of numbers tell the whole story.</p><p><strong>The jobs nobody counts.</strong> The World Economic Forum projects that by 2030, technological and structural change will displace roughly <strong>92 million jobs</strong> while creating about <strong>170 million</strong>, a net gain near <strong>78 million</strong>. That net figure is not a comfort, because, as the report itself stresses, those are different people doing different work. Around <strong>39% of current skills</strong> are expected to change by 2030, and <strong>85% of employers</strong> say they plan to reskill staff. Creation is happening. The question is whether anyone is pointed at it.</p><p><strong>The reason evaluation is now urgent.</strong> A 2026 enterprise survey found that <strong>88% of organizations</strong> experienced a confirmed or suspected AI agent security incident in the prior year. Prompt injection has moved from research curiosity to live exploitation, appearing in a majority of production deployments. Real examples from this year alone: a backdoored build of LiteLLM, the model gateway behind many popular agent frameworks, sat on PyPI for three hours and was downloaded roughly <strong>47,000 times</strong>. GrafanaGhost let hidden instructions pull enterprise data out through an image-rendering path. Multiple CVEs hit mainstream coding agents, letting a poisoned repository or allowlisted command deliver an attacker's payload.</p><p>Meanwhile, the WEF reports that only about <strong>14% of organizations</strong> say they have adequate AI security talent.</p><p><strong>What it pays.</strong> Full-time AI red teaming and evaluation roles in the US generally run <strong>$110,000 to $220,000</strong>, with senior and frontier-lab research roles reaching <strong>$280,000 and above</strong>. Contract work averages around <strong>$67/hour</strong>, with experienced specialists clearing <strong>$120 to $200/hour</strong>. Entry-level AI safety evaluator roles start in the <strong>$100,000 to $150,000</strong> range, and some firms hire at $60,000 to $70,000 with a bachelor's degree and genuine interest. Posted AI red team roles grew roughly <strong>55%</strong> across 2025 and 2026, pushed hard by the EU AI Act's August 2026 enforcement date, which requires red teaming for high-risk systems.</p><p>Put those together. A category of work that barely existed, demand rising on a regulatory deadline, a serious talent shortage, and an 88% incident rate in the field it protects.</p><p>That is not a side conversation about the future of work. That is the opening.</p><hr><h3><strong>๐ In the Wild</strong></h3><h3><strong>1. Building with AI is carving, not stacking</strong></h3><p>Mala started with a distinction that reframes the whole job.</p><p>Twenty years ago, software was <strong>additive.</strong> You started with nothing and built upward, feature by feature, until you had a product. Generative AI is <strong>reductive.</strong> She compares it to a master architect starting with an enormous slab of limestone and carving away until the form emerges. A large language model arrives as a vast, nebulous, general-purpose thing that can answer almost anything about almost anything. The work of building a real product is deciding what it should <em>not</em> do, and then proving it does not do it.</p><p>That carving happens two ways. <strong>AI safety research</strong> happens mostly inside the frontier labs and safety institutes, where there is enormous compute, deep specialist expertise, and direct access to the models. Almost nobody else can work at that level.</p><p>So a second discipline grew up around everyone else: <strong>AI evaluation.</strong> Taking what comes out of the box and asking whether it is fit for purpose in a specific context. Is it showing algorithmic bias? Is it hallucinating? Is it factually accurate, and is it factually <em>complete</em>, meaning does it give a half-answer that technically is not wrong but leaves out what the person actually needed?</p><p>Mala's marker of how fast this is moving: a year ago, saying "AI evaluation" to a general audience produced blank stares. Today a meaningful number of people know what it means, and more importantly, understand they have a role in it.</p><h3><strong>2. What a good evaluation actually looks like</strong></h3><p>Ask most people what an AI evaluator does and you get something vague about testing. Mala's definition is sharper, and it is the part worth stealing.</p><p>A good evaluation <strong>clearly states the problem statement being evaluated.</strong> You draw a hard boundary around what is in scope and what is out, then design a series of tests to see how the system performs inside that context. You can scope by vulnerability, exploit, or harm. You can run it with subject matter experts, or with lay users if you are about to scale a consumer tool to ten million people whose composition you cannot predict.</p><p>Then there are two modes:</p><p><strong>Adversarial testing</strong>, often called AI red teaming, a practice borrowed from cybersecurity. You are deliberately trying to break the system. If a product has a guardrail saying it must never answer finance questions, your job is to find the creative path around it.</p><p><strong>Non-adversarial testing</strong>, where you are not trying to break anything. You are trying to behave like a real user. You will find fewer dramatic vulnerabilities, but you will cover far more of the realistic scenarios where a system quietly fails someone.</p><p>Most organizations do neither, then call a meeting when something goes wrong in production.</p><h3><strong>3. The skills are not the ones you would guess</strong></h3><p>This is the part that should land for anyone wondering whether there is a door into this work.</p><p>Mala's answer for what makes a good adversarial tester: <strong>creativity, more than anything else.</strong> If you are a strong writer, if you speak multiple languages, if you can draw parallels between unrelated things, you can probably find ways to make a system do what it was told not to do. Technical depth in code or math helps when it is in scope, but it is not the entry requirement people assume.</p><p>For non-adversarial testing, the key skill is <strong>understanding the mindset of the user.</strong> Her teams set scenarios and ask testers to inhabit them: pretend you are this person, approaching this system with this need.</p><p>And the reason humans remain necessary is not sentimental. A model operates inside the parameters and data it was given. People carry lived experience, multiple languages, travel, exposure to other people, decades of memory. Mala made the contrast concrete: a model's context window is the last few turns of conversation, not forty years of being alive. That is a very different kind of context, and it has not been automated.</p><p>There is also a clear transfer path. Mala's own field, international development, has a discipline called monitoring, evaluation and learning. Practitioners who spent careers designing baselines and measuring whether a program worked are moving into AI evaluation, because the statistical methods carry over. Measuring human performance and measuring machine performance are not the same thing, but the methods rhyme.</p><p>If you have ever been the person who asks "what could go wrong here," you already have the instinct this work is built on.</p><h3><strong>4. Human judgment is infrastructure, not overhead</strong></h3><p>I put it to Mala directly: too many organizations treat human judgment as a budget line to be trimmed, when it may be the only thing standing between them and a serious failure.</p><p>Her answer was sharp. There are now papers showing that using a language model as the judge in an evaluation outperforms human annotators. But that conclusion rests on a hidden assumption: that <strong>better means consensus.</strong> Humans are not built to always agree, and that disagreement is a feature of a healthy society, not a defect in the measurement. Getting a hundred people to a single verdict is hard. Getting a machine to repeat the same verdict is trivial. Treating the second as superior quietly defines away the thing you wanted humans for.</p><p>Her verdict on humans as overhead: lazy and short-sighted. If you want to live in a world run entirely by machines, you can budget for that. If you believe AI should augment people rather than replace them, there is no argument to have.</p><p>She was equally honest about the real tension underneath, and it is one leaders are feeling right now. <strong>AI is far more expensive than the industry expected.</strong> Data centers, subsidized tokens, annotation, the sheer volume of data required to keep these systems current. That creates genuine trade-offs: a truly comprehensive evaluation of everything a model could do can run into the millions, and that same money might fund three years of your team. Nobody has a clean answer yet.</p><p>But the sequencing matters more than the budget. At Humane Intelligence, the problem space is <strong>always</strong> defined by humans first, then the evaluation is designed around it. The opposite approach, which Mala summarized with a laugh, is what most companies do: automate first, discover it broke, then throw humans at it with no thought given to the design. Oops.</p><p>Her favorite line on the economics of all this: the distance from <strong>zero to demo</strong> has gotten very small. The distance from <strong>demo to enterprise</strong>, or to anything sustainable, is still enormous.</p><h3><strong>5. Point it where the labor does not exist</strong></h3><p>The most human moment of the conversation was a story from Mala's first decade, working in tech for international development with the UN.</p><p>She was in Burundi, at the time the fifth poorest country in the world, trying to build a triage system for maternal mortality using SMS, about the most basic digital technology there is. The blocker was not the technology. It was that there were not enough doctors and nurses to do the triage at all. Even if more patients sent messages asking for help, nobody was there to read them. Training enough medical staff would have taken years and money the government did not have.</p><p>She said she wishes she had generative AI then, because that is the perfect use case. The labor did not exist. The systems to create the labor did not exist.</p><p>So her call to technology leaders is this: <strong>focus first on the places where the labor does not exist</strong>, rather than racing to automate what people are already doing. One creates capability that was never there. The other creates a fight over work that already belongs to someone.</p><p>That is the clearest definition of AI for good I have heard in a while.</p><hr><h3><strong>๐ฌ The Big Question</strong></h3><p><strong>If your company deployed an AI system tomorrow, who in your organization is responsible for proving it is safe, and do they have the authority to say no?</strong></p><p>And for anyone reading this wondering where they fit in the future of work:</p><p><strong>Are you waiting to find out which jobs survive, or are you looking at the jobs being created right now?</strong></p><p>Because the people who will do well in this transition are not necessarily the ones who build the models. They are the ones who can tell when a model is wrong, and explain why it matters.</p><hr><h3><strong>๐ง A Small Exercise</strong></h3><p>Mala's practical advice, split by who you are.</p><p><strong>If you are an individual wondering where to start:</strong></p><p></p><ol><li><p><strong>Separate upskilling from education.</strong> A stack of online certificates does not replace basic literacy and logical thinking. Build the foundation somewhere real, then layer the specialized training on top.</p></li><li><p><strong>Experiment where you already are.</strong> You do not have to become an evangelist. Be the person who first suggests it and brings it into your workplace. That is how Mala built a technology career on an undergraduate marketing degree and a master's in international affairs.</p></li><li><p><strong>Document everything.</strong> Put up a simple website. A portfolio, a few papers, a page of what you have done. She credits her personal site with a significant share of her speaking work, consultancies, and recruiter interest.</p></li></ol><p></p><p><strong>If you lead a team or a company:</strong></p><p></p><ol><li><p><strong>Define the problem space with humans first,</strong> before you design any evaluation. Bring in the people who will actually use the system, or who will bear the consequences of it. They will tell you which intersections matter, not just the obvious demographics.</p></li><li><p><strong>Decide your evaluation scope on purpose.</strong> Write down what the system must never do, then go try to make it do that. If nobody in your organization has run that test, you do not know what you have deployed.</p></li></ol><p></p><hr><h3><strong>๐๏ธ This Week on The Human Protocol Podcast</strong></h3><p><strong>The Job Nobody Is Counting: AI Evaluation and the Future of Work</strong> <strong>Guest:</strong> <a target="_blank" rel="noopener noreferrer nofollow" class="text-[#4FD1C5] underline hover:opacity-80 ember-view" href="https://www.linkedin.com/in/malakumar/"><strong>Mala Kumar</strong></a> - Executive Director of Humane Intelligence, novelist, and a global leader in technology for social good with experience across AI safety, open source, UX research, GitHub, the UN, and the World Health Organization</p><p>Mala is building the practices, people, and communities that make real AI evaluation possible. She is also refreshingly direct about cost, labor, and what the industry is getting wrong.</p><p>In this episode, you will learn:</p><p></p><ul><li><p>Why the AI jobs conversation is incomplete, and what regulation cannot solve on its own</p></li><li><p>Why building with generative AI is reductive, closer to carving stone than stacking bricks</p></li><li><p>What a good AI evaluation actually looks like, and the difference between adversarial and non-adversarial testing</p></li><li><p>The skills that make someone good at this work, starting with creativity</p></li><li><p>Why "LLM as judge" outperforms humans only if you believe consensus equals quality</p></li><li><p>Why human judgment belongs in the design from day one, not after the automation breaks</p></li><li><p>Where technology leaders should point AI first: at the labor that does not exist</p></li></ul><p></p><p>๐ง Available now on <a target="_self" rel="noopener noreferrer nofollow" class="text-[#4FD1C5] underline hover:opacity-80 foSeUQNNUAlOofCgTsBsMkWwskcxammaBWdAw " href="https://open.spotify.com/episode/48UQ9hOCFU5XjB5bODb2nt?si=h6DTQy0ESmC08fUzcXFOXw"><strong>Spotify</strong></a>, <a target="_self" rel="noopener noreferrer nofollow" class="text-[#4FD1C5] underline hover:opacity-80 foSeUQNNUAlOofCgTsBsMkWwskcxammaBWdAw " href="https://podcasts.apple.com/us/podcast/why-ai-is-way-more-expensive-than-companies-expected/id1843402023?i=1000792379292"><strong>Apple Podcasts</strong></a>, and <a target="_self" rel="noopener noreferrer nofollow" class="text-[#4FD1C5] underline hover:opacity-80 foSeUQNNUAlOofCgTsBsMkWwskcxammaBWdAw " href="https://youtu.be/_LVkvhgysq4?si=vI9eTjMcJ1e4hQVt"><strong>YouTube</strong></a>.</p><img class="max-w-full h-auto rounded-md my-4" src="https://media.licdn.com/dms/image/v2/D4D12AQFDYRtfMYhY_Q/article-inline_image-shrink_1000_1488/B4DaD9KBYOK0AI-/0/1790953658135?e=1792627200&v=beta&t=tPOO5pZGwNBmpcLFTxriCZKSw94pe5oSIoZngXFSmso" alt="Article content"><hr><h3><strong>๐งญ Thought to End the Week</strong></h3><p>I asked Mala if she is hopeful. Her answer stayed with me.</p><p>She said that during previous technology shifts, social media, cloud computing, the average person did not understand they had a role to play. They took what they were given. What is different now is that people are questioning, learning, organizing, and trying to build the skills to investigate these systems themselves. That participation is new, and it is real.</p><p>Then she added the part that ties it together: work to change the systems too. Build the society where people can experiment, make mistakes, and retrain without falling off a cliff. Because when people are motivated purely by survival, that changes what gets built and how.</p><p>Here is where I land. The future of work is not only automation. It is <strong>evaluation.</strong> The machines are going to keep getting more capable, and every gain in capability raises the value of the humans who can judge the output, catch the failure, and understand the consequence.</p><p>That is not a consolation prize for people whose jobs changed. It is one of the most important jobs of the next decade, and it is open right now.</p><p>Somebody has to check the machines.</p><p>It might as well be you.</p><p>The code may change. And it will.</p><p>But our humanity is the constant.</p><p>See you next week.</p><p><strong>Mykel</strong></p><hr><h3><strong>๐ Sources</strong></h3><p></p><ol><li><p>World Economic Forum, <em>Future of Jobs Report 2025</em>: 170 million jobs created and 92 million displaced by 2030, a net gain of 78 million; 39% of current skills expected to change; 85% of employers plan to reskill. <a target="_blank" rel="noopener noreferrer nofollow" class="text-[#4FD1C5] underline hover:opacity-80" href="https://www.weforum.org/press/2025/01/future-of-jobs-report-2025-78-million-new-job-opportunities-by-2030-but-urgent-upskilling-needed-to-prepare-workforces/"><strong>https://www.weforum.org/press/2025/01/future-of-jobs-report-2025-78-million-new-job-opportunities-by-2030-but-urgent-upskilling-needed-to-prepare-workforces/</strong></a></p></li><li><p>Help Net Security and OWASP GenAI Security Project (2026): prompt injection as the dominant agentic AI failure mode; the LiteLLM supply chain backdoor downloaded roughly 47,000 times in three hours; CVEs against major coding agents. <a target="_blank" rel="noopener noreferrer nofollow" class="text-[#4FD1C5] underline hover:opacity-80" href="https://www.helpnetsecurity.com/2026/06/11/owasp-prompt-injection-ai-security-failures/"><strong>https://www.helpnetsecurity.com/2026/06/11/owasp-prompt-injection-ai-security-failures/</strong></a></p></li><li><p>OWASP GenAI Exploit Round-up Report Q1 2026: GrafanaGhost indirect prompt injection and data exfiltration; Flowise CVE-2025-59528 active exploitation. <a target="_blank" rel="noopener noreferrer nofollow" class="text-[#4FD1C5] underline hover:opacity-80" href="https://genai.owasp.org/2026/04/14/owasp-genai-exploit-round-up-report-q1-2026/"><strong>https://genai.owasp.org/2026/04/14/owasp-genai-exploit-round-up-report-q1-2026/</strong></a></p></li><li><p>2026 enterprise survey cited in agentic AI security research: 88% of organizations reported a confirmed or suspected AI agent security incident in the prior year. <a target="_blank" rel="noopener noreferrer nofollow" class="text-[#4FD1C5] underline hover:opacity-80" href="https://arxiv.org/pdf/2607.05518"><strong>https://arxiv.org/pdf/2607.05518</strong></a></p></li><li><p>The Interview Guys, <em>Best AI Red Teaming Jobs in 2026</em>: salary bands from $110,000 to $300,000+, contractor rates to $200/hour, and EU AI Act driven demand. <a target="_blank" rel="noopener noreferrer nofollow" class="text-[#4FD1C5] underline hover:opacity-80" href="https://blog.theinterviewguys.com/best-ai-red-teaming-job/"><strong>https://blog.theinterviewguys.com/best-ai-red-teaming-job/</strong></a></p></li><li><p>ZipRecruiter (September 2026): average US AI red teamer pay of $67.60/hour, roughly $140,600 annualized. <a target="_blank" rel="noopener noreferrer nofollow" class="text-[#4FD1C5] underline hover:opacity-80" href="https://www.ziprecruiter.com/Jobs/Ai-Red-Teamer"><strong>https://www.ziprecruiter.com/Jobs/Ai-Red-Teamer</strong></a></p></li><li><p>TechJack Solutions AI career guide, citing WEF 2025: only 14% of organizations report adequate AI security talent; entry pathways and salary ranges. <a target="_blank" rel="noopener noreferrer nofollow" class="text-[#4FD1C5] underline hover:opacity-80" href="https://techjacksolutions.com/careers/ai-careers/ai-red-teamer/"><strong>https://techjacksolutions.com/careers/ai-careers/ai-red-teamer/</strong></a></p></li><li><p>Humane Intelligence: human-centered AI evaluation, red teaming practice, and LinkedIn Learning course. <a target="_blank" rel="noopener noreferrer nofollow" class="text-[#4FD1C5] underline hover:opacity-80" href="https://humane-intelligence.org"><strong>https://humane-intelligence.org</strong></a></p></li></ol><p></p>