The Hiring System as an RLHF Agent

arhngl1 pts0 comments

The Hiring System as an RLHF Agent - arhngl

arhngl

SubscribeSign in

The Hiring System as an RLHF Agent<br>Why Good Specialists Can't Find Jobs for Months

arhngl<br>Aug 19, 2026

Share

Lately, LinkedIn has seen more and more posts from developers and other IT professionals who have been looking for a job for months, and sometimes even years.<br>And this is not a story about a random rejection at the screening stage. People go through several stages of interviews, live coding, technical discussions, talk to a hiring manager, and sometimes make it almost to the very end – only to be met with silence.<br>One such case could be written off as a bad process, an overloaded recruiter, or an unlucky candidate. But when the exact same pattern repeats at scale, another question arises:<br>What if the problem is no longer with individual people, but with the system itself?

It seems to me that the analogy with RLHF fits this surprisingly well.<br>Not in the sense that HR literally works and learns like an LLM. But in the sense that a system can gradually learn to optimize for its own local reward signals while moving further and further away from its original task.<br>The basic task of hiring seems simple enough:<br>Find a person who will solve the problem the business needs solved.

But in practice, the system might be optimized for something entirely different:<br>Run the process in such a way that the decision looks safe and well-justified to all participants.

And these are fundamentally different functions.

Reward Function Shift: From Results to Safety

In an ideal model, the reward looks roughly like this:<br>hired a strong engineer → they solve problems → the business gets results.

But inside a large company, every participant in the process has their own reward. The recruiter needs to fill the vacancy and avoid making an obvious mistake. The hiring manager wants a candidate who is easy to defend to leadership. Legal doesn’t want risky wording or admissions. The team wants a person who will predictably fit into the existing structure.<br>This naturally gives rise to CYA (Cover Your Ass).<br>All else being equal, it is safer for the system to choose a candidate with a “glossy” resume: a clear track record, 5+ years in a single technology, a major brand in their experience, and an exact keyword match (or to hire no one at all, as paradoxical as that may sound). Even if they are not the best candidate.<br>If such a hire turns out to be a mistake, it is easy to justify:<br>They had an exemplary pedigree and the right stack, no one could have foreseen the error.

But if you hire an unconventional person with a short, choppy career path, several stack changes, and a lot of experimental experience, explaining such a decision is much harder.<br>And yes, currently this applies not only to unconventional specialists. The market is so overloaded with formalities that even the “right” candidate gets stuck in the same pipeline.<br>This creates a strange situation:<br>The system minimizes not the probability of a bad hire, but the risk of looking wrong at the moment of hiring.

AI as a Catalyst: New Context, Old Filters

Over the past two years, tools like Claude Code, Cursor, and Codex have become prominent enough to change the context. What matters is not whether they are replacing developers right now, but how the very fact of their existence changes the business’s expectations at the entry point of the hiring pipeline.<br>At the business level, often subconsciously, a candidate starts to be perceived not as a scarce specialist, but as a resource potentially made cheaper by a tool. And when the business receives the signal that “part of the work can be automated,” it increasingly needs a different type of person – someone who can:<br>quickly learn new tools;

work with AI as a partner;

figure out poorly defined tasks;

see systemic problems;

relearn faster than the stack changes.

The value of a person shifts from knowing a specific framework to the ability to understand a new task, choose a tool, and solve a problem that nobody knew how to solve yesterday.<br>The paradox is that the market already wants more adaptable people, but the hiring system, being inert, continues to look for those who best fit the historical template. As a result, the new context does not change the selection mechanics itself – it merely adds another layer of noise to a system where safety already dominates.

Epistemic Gap: The Company Says One Thing, the Filter Looks for Another

A large part of the hiring pipeline is built on pattern matching: company name, years of experience, specific technology, job title, sequential career path.<br>In a stable industry, this works quite well: past patterns can indeed be predictors of future success. But during a technological shift, this principle breaks down.<br>An engineer might, over a few years, switch two or three stacks, leave a corporation for a startup, experiment with AI, or transition from backend to ML. To a human, this could be a story of high...

hiring system candidate from years person

Related Articles