Building the enterprise environment for agentic AI


  1. Agentic AI is a larger systems problem, not just one of inference.
  2. The majority of existing agentic AI harnesses are limited and do not measure overall system performance.
  3. Plan capacity is done using agents per virtual CPU (vCPU) density, not agent count.
  4. Monitor agent task latency, not just average CPU utilization.
  5. Default to scale-out for systems hosting agents. Reserve scale-up for workloads with heavier per-agent compute or architectural constraints.

Beyond inference: Agents as workflow automation

Agentic AI is more than LLM inference. Its enterprise value depends on the full system, task orchestration, data access, tool execution, latency management, governance, and scalable infrastructure. An agent is a goal-driven automated enterprise workflow process: It plans a multi-step task, calls tools, reads results, and retries when something fails. Enterprise agents are therefore not just an inference problem; they are a systems problem.

Defining what good looks like

Most agentic AI metrics focus on evaluating the LLM used. Platform teams also need to know how long the tasks take, how many agents a fleet can support, what users experience at the end of the execution process, and how costs change as more agents work simultaneously.

A more useful enterprise view looks at six metrics:

  1. Task success rate
  2. Cost per task
  3. Time per task
  4. Task throughput
  5. Agent density (agents per vCPU)
  6. Latency

Together, these answer the questions enterprise AI operators care about: Is the system performing as expected? How many agents can the system sustain? How should it scale to support more agents?

Building on solid foundations

To gain a deeper insight into agentic AI workload performance, Intel extended Terminal-Bench, an open source benchmarking harness for evaluating AI agents with profiling, telemetry, and replay capabilities. This made it possible to understand where the agents spent time beyond LLM inference.

The benchmark extension used a deterministic record-replay of LLM responses to separate agent performance from LLM variability. LLM responses were recorded once and replayed identically across runs, reducing run-to-run variance and creating a more reliable basis for comparison.

The Terminal-Bench task mix used was intentionally broad. It included compilation, testing, database operations, Boolean logic, interpretation, ray tracing, compression, linear algebra, video transcoding, and machine learning training. That wide variety made the findings more relevant to real enterprise environments.



Source link

  • Related Posts

    New Firefighting Technologies Could Help Battle Blazes Like Those in France and Spain

    Spain declared a national emergency in the community of Madrid and the province of Ávila due to the spread of intense wildfires. With fires still burning out of control, this…

    Satya Nadella says companies that trust one AI for everything may not survive

    On Sunday, Microsoft CEO Satya Nadella doubled down on the shocking warning he issued earlier this month to businesses that use AI, taking it a step further this time. Companies…

    Leave a Reply

    Your email address will not be published. Required fields are marked *

    You Missed

    'Onslaught' Trailer 2

    'Onslaught' Trailer 2

    D4vd murder case: Singer to stand trial on charges over 14-year-old girl’s death

    D4vd murder case: Singer to stand trial on charges over 14-year-old girl’s death

    New Firefighting Technologies Could Help Battle Blazes Like Those in France and Spain

    New Firefighting Technologies Could Help Battle Blazes Like Those in France and Spain

    The colorful, funky new arrivals I’ve had my eye on all month.

    The colorful, funky new arrivals I’ve had my eye on all month.

    Alberta to begin allowing self-referred medical testing as of Friday

    Alberta to begin allowing self-referred medical testing as of Friday

    A Japanese town wrestles with identity after protests over its first mosque

    A Japanese town wrestles with identity after protests over its first mosque