Aug 26 2026
Data Center

Intel’s Xeon 6 Portfolio Broadens the Scope of the CPU in Enterprise AI

As businesses advance artificial intelligence projects, CPUs provide IT leaders with options for allocating compute resources based on the specific performance, scale and efficiency demands of each workload.

Over the past few years, discussions regarding enterprise artificial intelligence infrastructure have focused primarily on GPUs. The reason is clear: GPUs provide the massive amounts of parallelism required to develop and deploy large language models, as well as some of the most computationally intense AI workloads.

However, as organizations transition beyond the experimental phase of AI adoption and begin integrating AI throughout various business functions, the complexity of developing and managing AI infrastructure grows. Many organizations are now using pretrained models instead of developing them from the ground up, which shifts a larger portion of the AI lifecycle toward inference.

Additionally, agentic AI increases the demand to integrate AI models with applications, data and systems currently running on traditional, CPU-based infrastructure.

According to David Bartley, an Intel channel account manager, “CPUs have been running AI workloads for decades.” He added, “The narrative only shifted when generative AI showed up, and suddenly everything was seen through the eyes of a GPU.”

DISCOVER: Intel technologies can help support your AI investments.

The timing wasn't an accident. Training large language models demands enormous parallelism, exactly what GPUs are built for, so as generative AI captured attention, the GPU became shorthand for AI infrastructure itself.

With its new Xeon 6 portfolio, Intel is providing organizations with a broad family of socket-compatible processors specifically designed to accommodate diverse requirements across the entire data center, from high-performance AI inference to densely packed, energy-efficient cloud computing.

CPUs Play a Critical Role in Enterprise AI

GPUs will remain instrumental for many AI workloads, especially those involving large-scale model training. But training is not necessarily the only — or even the most significant — aspect of AI activities for many organizations.

Rather than creating their own proprietary foundational models, many enterprises are choosing to use pretrained models as a starting point for their AI initiatives, then fine-tuning those models to suit their specific business requirements. Once those models go live in production environments, inference typically constitutes the largest part of their operational lifecycle. Many of these applications use smaller, specialized models that are accessed by a relatively small number of employees.

For instance, a legal department may use a retrieval-augmented generation application to analyze large volumes of contract documents and answer questions related to their content. Similarly, HR departments and customer service centers may use similar types of targeted AI tools. Such workloads present opportunities for organizations to leverage their existing Xeon infrastructure as opposed to automatically investing in dedicated GPU systems, Bartley says: “Most of our customers start with internal use cases, then expand to customer-facing ones as they get comfortable.”

Click the banner below to learn how to turn complexity into a business advantage.

 

Intel’s Xeon 6 portfolio includes Intel Advanced Matrix Extensions, or AMX, an integrated accelerator designed to enhance performance for AI-related matrix operations. For certain smaller-model inference workloads, AMX enables organizations to deploy AI applications using CPU-based infrastructure without requiring additional GPUs.

In addition, the emergence of agentic AI may further increase the role played by the CPU. As AI agents evolve beyond merely generating responses and begin taking actions across enterprise applications, more of the work becomes orchestration: calling tools and APIs, retrieving data, executing functions, and moving results between systems. That glue is CPU work. And because the enterprise applications and the data these agents rely on already run predominantly on x86 infrastructure, the CPU is a natural place to host that integration layer, running natively alongside those systems rather than translating across architectures.

GET THE DETAILS: Build a data infrastructure that supports AI initiatives.

How Xeon 6 Supports Workload-Specific AI Infrastructure

The growing need for workload-specific infrastructure is reflected in Intel’s Xeon 6 portfolio. Rather than using one type of processor architecture for every use case, organizations can select processors tailored to meet varying priorities while retaining socket compatibility among members of the portfolio. This adaptability may prove increasingly beneficial as organizations strive to balance AI performance with cost-effectiveness, energy consumption, space utilization and overall operational efficiency.

When planning for AI infrastructure, IT leaders should always consider the workload first:

  • Which model will be employed by the organization?
  • How many users will require access to it?
  • What are the required response times?
  • Is this application intended for internal use or external customers?

These questions will assist in determining whether a CPU-based solution is adequate, whether GPUs are required, or whether a hybrid approach combining both CPU- and GPU-based solutions is optimal.

“We’re not advocating that you use Intel CPUs for every possible use case,” Bartley says. “Large-scale models, applications serving very large numbers of concurrent users, and applications that demand extremely low latency all continue to see significant advantages from GPUs. But the small language models that make up the majority of what our customers deploy serve a relatively modest number of concurrent users, and that workload sits comfortably within what a Xeon with AMX can handle. For those, this is an architecture worth investigating.”

Finally, financial considerations also play a role in making decisions regarding AI infrastructure. As organizations continue to rely more heavily on cloud-based, API-driven AI services, token-based costs associated with those services are becoming an increasing source of concern.

Operating appropriately sized workloads on existing or new Xeon infrastructure may provide some organizations with greater control over long-term AI operating expenses.

Brought to you by:

hapabapa/Getty Images
Close

New Research from CDW Explores AI and Cybersecurity

Learn how AI is helping IT teams manage risk and improve resilience.