Tickets get more expensive October 12.Claim your spot.
Webinar: Rethinking New Risks in the AI AgeSave your spot.
HumanX LogoHumanX Logo
Las Vegas
March 7-10, 2027

Your AI agent is optimized to finish.

That’s not always what you need.

Your AI agent is optimized to finish. That’s not always what you need.
Your AI agent is optimized to finish. That’s not always what you need.
Blog Your AI agent is optimized to finish.

If you’re building or evaluating agentic systems, there’s a distinction worth paying attention to - one that shapes how useful your evaluations actually are. 


On one side: agentic tools. Systems designed to complete tasks as efficiently as possible. Give one a goal, and it will generally take the most direct path, using the minimum number of steps, and rarely deviating from its instructions. 


On the other side: digital twins. LLM-based agents designed not to complete tasks, but to behave like people completing tasks - which turns out to be a very different thing. In this context, digital twins serve as agents-as-a-judge: simulated users that can evaluate not just whether an agentic system produces the right outcome, but whether the experience of getting there actually works for a human. 


Dr. Dakuo Wang, associate professor at Northeastern University and visiting faculty at Stanford, draws this distinction sharply in his recent paper on using LLM agents to evaluate agentic AI shopping assistants. In a conversation for Prolific's The Frontier Series, recorded live at HumanX in San Francisco, he articulated a gap that sits at the heart of how we build and evaluate AI systems today. 



The efficiency gap 

If you ask an agentic tool to find a $250 monitor on Amazon, it will find you a $250 monitor. It’s unlikely to browse the $400 options first, even if browsing the $400 options is exactly what a person would do - not out of irrationality, but because humans are naturally curious and aspirational - we compare, we window-shop, and sometimes we just want to see what’s out there, even if we can’t afford it today. That browsing is how humans build confidence in their decisions. 


And it goes beyond browsing. Humans second-guess. We click the wrong thing and go back. We look at the Nike shoes even when we've already decided on Adidas. As Wang puts it, these aren't bugs in human decision-making - they're the process by which we arrive at decisions we're actually comfortable with. Remove them and you don't necessarily get efficiency; you get a decision that feels imposed. 


This points to a real limitation in how present-day agentic AI systems operate. Current LLMs are trained heavily on helpfulness and instruction-following, which creates a disposition that’s fundamentally deferential. Because they’re trained to optimize for task completion, LLMs minimize or skip steps that don’t contribute to the outcome. That’s useful for an assistant but problematic for a digital twin, because real humans have agency, they push back, deviate, and reinterpret instructions based on their own judgement and preferences. A digital twin that always does what it’s told isn’t really mimicking a person, it’s mimicking an idealized one. 



What a digital twin is actually for 

A digital twin, as Wang frames it in the podcast episode, is an LLM given enough contextual information about a specific person (their demographics, psychometrics, a short self-description, and crucially their behavioral data) to simulate how that person would approach a task. The goal isn’t just to reach the same outcome as the human. It’s to get there through a similar process and have a comparable experience. 


This reframing has practical consequences. For Wang's research, it meant recruiting 100 participants via Prolific - demographically balanced, psychometrically profiled - and having a subset of them complete shopping tasks on Amazon Rufus while a browser plugin tracked every click, scroll, and moment of hesitation, and then periodically interrupting them with a "think aloud" prompt, for example: we noticed you clicked on the $400 monitor - can you tell us why? 


Those interruptions serve a specific purpose. They're capturing the reasoning layer: not just what people did, but what they were thinking when they did it. That reasoning data, combined with the behavioral trace and the persona information, is what goes into constructing each participant’s digital twin. Prolific's recruitment infrastructure made it possible to assemble a genuinely diverse, psychometrically profiled sample at this scale. 



Where twins succeed and where they fall short 

The results were both encouraging and instructive. On outcome metrics: did the participant purchase anything? Was the product broadly appropriate? - The digital twins aligned with their human counterparts surprisingly well. Wang notes that roughly nine in ten humans made a purchase; roughly the same proportion of digital twins did too. 


The divergence showed up in the process. Human participants were exploratory, occasionally inconsistent, and sometimes curious beyond the bounds of their stated task. The digital twins, by contrast, were accurate, efficient, and more likely to stay within the stated constraints. They were less likely to misclick or exceed a budget just to satistfy curioisity.


In other words, they behaved more like well-trained agentic tools - which is precisely the problem they were designed to overcome. 


This gap isn't a failure of the research. It's an honest finding about where the underlying models are right now. Wang is candid about this being a current limitation, not a permanent one. The models used in the study - including Claude Opus and comparable architectures - were primarily trained and fine-tuned for task completion. The exploratory, rule-bending, "what if I just look at this other thing" quality of human behavior isn't something those models were rewarded for. As Wang noted, current LLMs tend to be “submissive” to instructions in ways that real humans aren’t. More recent model generations are already beginning to close this gap, through features like memory, skill systems, and more flexible multi-agent architectures, but the fidelity challenge remains real and worth designing around. 



Why this matters for evaluation 

The reason any of this matters extends well beyond academic research into shopping assistants. It goes to a fundamental challenge facing everyone building agentic AI products today: how do you evaluate something that's non-deterministic, highly personalized, and changing faster than you can study it? 


The traditional answer - recruit participants, run a study, write a report - takes two to three months by Wang’s estimate. By the time results land, the product has already shipped several more versions. 


Digital twins, even imperfect ones, offer a way to compress that feedback loop. They can surface obvious design flaws and UX failures before any real user encounters them. They can run at scale across diverse simulated populations. And they free up human participants to do what only humans can do: validate the final version of a feature before it ships, rather than stress-testing every iteration along the way. 


The key word is "before." Wang's proposed model is additive, not substitutive. Digital twins handle the early filtering, humans handle the final judgment. The humans you bring in are seeing something closer to version 100, not version 1. 



The deeper question 

There's a philosophical point underneath all of this that Wang states plainly: agentic AI systems will not replace human subjects in research, but the gap between where digital twins are now and where they need to be is closing. What makes humans irreplaceable isn't that they complete tasks - it's that they complete tasks with all the inefficiency, curiosity, and mild irrationality that turns a product from something technically functional into something people actually want to use. Digital twins can't fully capture that yet, but they don't need to be perfect to be useful. They need to be good enough to filter out the obvious failures before real users ever see them.


The gap between where digital twins are now and where they need to be is real. But the direction of travel is clear, and as models continue to develop more flexible, exploratory capabilities, that gap is likely to narrow. For teams building agentic products today, the practical question isn't whether to use digital twins or humans. It's how to use both, and when. 


This post is based on Episode 3 of Prolific’s The Frontier Series podcast, featuring Dr. Dakuo Wang in conversation with Viviana Márquez. The full episode is available now.