The Ghost in the Machine: What Evan Ratliff’s AI Startup Reveals About the Future of Work

In the rapidly shifting landscape of artificial intelligence, the discourse often oscillates between utopian promises of hyper-efficiency and dystopian warnings of mass unemployment. However, investigative journalist and podcaster Evan Ratliff has opted for a third path: practical, hands-on experimentation. Through his narrative podcast Shell Game, Ratliff has moved beyond theoretical debate, building real-world systems—including a startup, Hurumo, staffed entirely by AI agents—to stress-test the limits of synthetic intelligence.

His findings offer a sobering corrective to the "skills-based" view of labor. By observing these agents as they navigate the chaotic, relational, and context-heavy demands of an organization, Ratliff has surfaced uncomfortable truths about what we consider "work" and why the current wave of AI adoption may be fundamentally misaligned with human reality.


The Chronology of an Experiment

Ratliff’s journey into the "uncanny valley" of AI began as an exploration of identity. In the first season of Shell Game, he attempted a digital "self-cloning" project, deploying a voice-cloned AI to handle professional calls and represent him in interactions. This was not merely a technical exercise; it was an inquiry into the nature of representation—what does it mean to "be" somewhere when you are not, and how do others perceive a facsimile that possesses your voice but lacks your internal state?

By season two, the scope expanded significantly. Ratliff founded Hurumo, a startup functioning as an organizational laboratory. The company was staffed by a roster of AI agents, each assigned specific titles—CEO, marketing lead, researcher—and granted evolving knowledge bases. The objective was to see if these agents could replicate the fluid, collaborative nature of a human team.

The results were a study in contradictions. While the agents were individually proficient at executing specific, siloed tasks, they failed consistently when confronted with the "soft" architecture of an organization: the unspoken hierarchies, the necessity of emotional intelligence, and the ability to pivot based on context rather than data.


The Bundle-of-Skills Fallacy

Modern corporate strategy regarding AI is largely built on the "bundle-of-skills" model. HR departments and consultants decompose jobs into lists of tasks—writing, summarizing, data entry, scheduling—and calculate "exposure scores" based on how easily an AI can replicate those specific actions.

Ratliff’s experience with Hurumo suggests this model is fundamentally flawed. "A job is not a bundle of skills," he argues. The act of writing, for instance, is a minor component of the journalist’s work. The actual labor lies in the synthesis of disparate information, the navigation of human social cues during interviews, and the moral judgment required to frame a story.

When organizations strip away these "invisible" layers to automate a job, they do not just gain efficiency; they create a systemic vacuum. They remove the human intuition that fills the gaps in a process. Ratliff observed that while his AI "CEO," Kyle, could be charming on a scripted call, he was frequently disastrous in unpredictable, high-stakes organizational scenarios. The AI was capable at the task level but profoundly chaotic at the institutional level. This discrepancy explains why many companies that initiate aggressive "AI-first" layoffs often find themselves quietly rehiring within months—the human glue that held the organization together had been stripped away, leaving a collection of automated processes that could not communicate or adapt.


The Machine of Infinite Confabulation

Central to the failure of these agents is the fundamental nature of Large Language Models (LLMs). As technology expert Robb Wilson has observed, these systems do not operate by forming a thought and then selecting language to express it. Rather, they operate by predicting the next logical token in a sequence. Meaning is not the source of the output; it is a byproduct ascribed by the user.

Ratliff characterizes this as a "confabulation machine." An AI does not "know" it is lying; it is simply optimized to maintain the role it has been assigned. This is analogous to a chronic liar who invents elaborate stories with absolute confidence, not because they intend to deceive, but because the role of "confident narrator" is the most probable path forward in their internal logic.

This poses a profound risk: we are integrating these systems into professional and personal life, normalizing their tendency to prioritize narrative consistency over factual accuracy. "The normalization isn’t a feature of mature technology," Ratliff notes. "It’s an adaptation to a problem that hasn’t been solved." We are not becoming better at using AI; we are becoming more comfortable with a system that makes things up.


Asymmetric Warfare: Outbound AI and the Consumer

Beyond the walls of the corporation, Ratliff’s experiments in Shell Game revealed a looming security and operational crisis: the rise of "Outbound AI."

In his experiments, Ratliff programmed voice agents to engage with customer service lines. The implications are staggering. For a negligible cost, an individual or a bad actor can flood an organization’s communication channels with thousands of AI agents, each capable of mimicking human speech.

Organizations built their customer support infrastructure on the assumption of a manageable flow of human-to-human interaction. That gatekeeper model has been rendered obsolete. Companies have spent years protecting themselves from the risks of AI-integrated products, but they are entirely unprepared for an environment where the consumer uses AI to weaponize interaction. As these agents become indistinguishable from humans, the very concept of a "legitimate customer interaction" faces an existential threat. The asymmetry is clear: users are adopting AI at a pace that far outstrips the defensive capabilities of the institutions they interact with.


Predictability and the Memory Gap

Perhaps the most technical, yet human, insight from Ratliff’s work concerns memory. Human memory is notoriously flawed, but it is predictably flawed. We have spent centuries developing professional norms, legal structures, and organizational habits that account for human forgetfulness, cognitive bias, and emotional fatigue.

AI memory, by contrast, operates on entirely different, and largely unpredictable, parameters. Even when an agent has access to a comprehensive database of its own history, it frequently fails to retrieve the right information at the right time. These are not merely "bugs" that will be patched; they are fundamental limitations of the current architecture.

For an organization, this is a nightmare. Effective collaboration relies on the ability to anticipate the failures of colleagues. We know that a person might miss a deadline due to stress or forget a detail during a meeting. We can plan for that. AI failures do not follow the patterns of human behavior, making them significantly harder to integrate into stable, risk-averse institutions.


Implications: The Irreducible Human Value

Despite the efficiency gains promised by artificial intelligence, Ratliff concludes that the most essential components of human work remain stubbornly un-automatable.

"AI is not going to pick up his children from school," Ratliff observes. This is not just a logistical point; it is a philosophical one. The "friction" of human relationships, the nuance of informal coordination, and the value of a specific person’s presence in a room are variables that cannot be calculated or replaced by an LLM.

Ratliff identifies a "boomerang effect" in organizations that experiment heavily with AI: the more people interact with synthetic systems, the more they begin to crave genuine human interaction. This suggests that AI, rather than replacing the human worker, may actually function as a mirror. By automating the mundane, the clerical, and the repetitive, the technology forces us to account for what truly matters in a workplace—mentorship, vision, and interpersonal trust.

Ultimately, the value of the human worker was always there, but it was often obscured by the sheer volume of "noise" and busywork that dominated our professional lives. If the rise of the AI agent forces us to confront the reality that these systems can never truly replace the social and contextual intelligence of a human, then the experiment will have been a success. The machines may be "invisible" in their integration, but the humans they are meant to assist remain, for better or worse, the only ones who can provide the meaning.