Skip to content

Hiring Agents Is the Easy Part

Authors
Alana Levin Caleb Shack
Tags
AI

Agents are becoming a much more ingrained part of how we do work. They expand the capabilities, capacity, and agency of each person within an organization. Many AI tools today focus on how to augment human workflows; we’re excited about much more work shifting toward agents that automate workloads, especially as domain-specific models lead to novel frontiers.

As we think about how to get from the status quo (largely augmentation) to the future (agents as automators), one analogy we’ve found useful is framing agents as akin to employees. With employees, there are tried-and-true processes for hiring, onboarding, and promoting people. It looks something like: 

  • Screening: testing whether the person is capable of doing the job.
  • Onboarding: integrating the new hires into existing workflows and org structures, setting up access (controls) for internal tools, and broadly providing company-specific context. 
  • Ongoing performance reviews: assessing whether the work produced is actually good and providing feedback on how to improve.

The same can and will be true for agents. And to be clear, parts of the agent hiring infrastructure are already well underway in being built out. Offline evals are good at assessing whether a model or agent can complete a specific workflow (step #1, aka screening). Infrastructure for step #2 might take the form of something like Glean or Cognee – shared substrates that let agents easily access an organization’s tools and leverage company knowledge graphs. 

Step #3 – assessing whether the work is “good” and providing feedback on how to make it better – is where we’re most interested. Some of the questions top-of-mind for us include:

  • How is “good” defined (and who defines it)? The first set of automators has focused on areas with easily verifiable work – i.e., was the job completed, yes/no? But in the future, it will have more nuance as quality verification gets more complex. The criteria may start to look like: was the job completed efficiently? Then, was the job completed to a company’s tacit standards (e.g., consistent with the intuitive, experiential knowledge within an organization)? We’re especially interested in work that has historically been hard to verify but is becoming easier as models improve.
  • How does “good” compound across agent lifecycles? Today, feedback on agents in production (“online evals”) largely guides agentic systems to fitted tasks. The loop looks like: a human logs an agent’s success or failure, labels it, scores it, and that feedback directs the agent toward improvement. But the eval lives as an artifact in an external, human-maintained database. Is there a world in which feedback can accumulate within the agent itself and more autonomously compound knowledge across the org? In other words, what would it take for agents to begin directing their own learning loops?
  • When an agent receives company-specific feedback, who owns that data? Palantir calls this “Sovereign AI”; we think of it as ownership as a keystone of the architecture (a key pillar in our autonomy thesis). 
  • Does every company need its own eval layer? When does a company need to customize an agent to its specific processes vs. in what scenarios should we expect an agent to work out of the box?
  • If an agent within an organization goes rogue, where does the liability fall? Questions regarding internal permissions, liability, and access controls are core to building hardened autonomous systems. Our in-house attorney/investors @dbarabander and @sabina_beleuz also admittedly ensure we spend sufficient reps thinking about how founders can avoid liability sinkholes 🙂 

The world 3-5 years from now will have a lot more agents in it than exist today. How we get from here to there has a lot of fuzzy questions – ones that we’ve encountered firsthand as we’ve built out our own agent systems, as we chat with portfolio companies about their incorporation of agents, and as some of our investments (e.g. Prism) tackle every day. 

It’s an area we’re actively excited about, as both investors and AI-native users, and would love to chat with people thinking about these questions.

View Similar

AI Article