What Is HPC (High-Performance Computing)?

Related problems: Simulations or models take days to run on our servers; Engineers waiting in line for limited compute capacity; Deciding whether to buy a compute cluster or rent capacity in the cloud; Our colocation space can't power or cool the hardware we need

High-performance computing (HPC) is the use of unusually powerful or specialized computing resources to run computationally intensive workloads. Parallel processing is the most common way to get that performance, but it is not a requirement: a demanding job on a single high-end node or workstation can also be HPC. The most common architecture is a cluster: many powerful servers, often with specialized processors such as GPUs, tied together by fast networking and fed by high-throughput storage, so a job that would take weeks on an ordinary server can finish in hours. Large shared-memory systems, in which many processors share one very large memory, are also HPC systems. HPC is used for engineering simulation, scientific research, financial modeling, rendering and, increasingly, AI training, and it can run in an organization’s own data center, in colocation, on hosted bare metal or in the public cloud.

At a glance

  • Work is usually split into pieces that run at the same time, most commonly across a cluster of compute servers called nodes, though large shared-memory systems and single high-end nodes are also used.
  • Low-latency, high-bandwidth networking between nodes is often as important as the processors themselves.
  • Parallel storage systems feed data to many nodes at once without becoming a bottleneck.
  • HPC hardware is dense and power-hungry, so facility power and cooling are often the limiting factor.
  • Can be owned, hosted, rented as bare metal or run in the cloud, with very different cost profiles.

What problem it solves

Some computing problems are simply too big for an ordinary server: simulating airflow over a vehicle, modeling how a drug molecule behaves, running thousands of financial risk scenarios overnight, rendering a film or training a large machine learning model. On ordinary infrastructure these jobs take too long to be useful, or cannot run at all because the data does not fit in an ordinary server’s memory.

HPC applies far more computing power than an ordinary server, usually by breaking the problem into parts and running them in parallel, most often across many servers sharing data over a fast network, or across the many processors of a large shared-memory system. The result is answers in hours instead of weeks, more design iterations, larger models and the ability to tackle problems that would otherwise be out of reach. For businesses, faster turnaround on simulation or analysis often translates directly into faster product development or better decisions.

How it works

Compute nodes. The descriptions below cover the cluster, the most common HPC architecture. An HPC cluster contains many servers with high core counts, large memory and, for many workloads, GPUs or other accelerators. Nodes are typically identical so work can be spread evenly.

Interconnect. Nodes exchange data constantly during a job, so the network between them is critical. HPC clusters often use specialized low-latency interconnects or high-speed Ethernet tuned for this traffic. Latency between nodes can limit how well a job scales.

Storage. Parallel file systems spread data across many storage servers so many nodes can read and write at once. Many clusters pair a fast “scratch” tier for active jobs with cheaper storage for results and archives.

Scheduler and software. A job scheduler queues work from many users and assigns it to available nodes. Applications use parallel programming libraries to split work across processors and nodes. Many industry simulation packages are licensed per core or per job, which can drive cost as much as hardware.

Facilities. Dense HPC racks can draw far more power than general-purpose racks, which drives requirements for high rack power density, advanced cooling such as liquid cooling and power delivery that many older facilities cannot support.

For hosting options, see our bare metal and colocation pages.

When it matters for buyers

  • Buy, host or rent. Owning hardware in your facility or in colocation can be cheaper for steady, heavy use; cloud and GPU as a service suit bursts and avoid capital spend. Model both on real usage.
  • Facility limits. Before buying hardware, confirm the data center can deliver the power and cooling per rack it needs.
  • Data gravity. Large input and result data sets are slow and costly to move; place compute near the data or plan transfer and egress costs.
  • Software licensing. Per-core or per-node licenses can make some deployment models much more expensive than others.
  • Refresh planning. Processor and GPU generations change quickly, and older clusters may become uneconomic before they wear out.

Questions to ask vendors

  • What processors, accelerators and interconnect do you offer, and how are nodes connected to each other?
  • What parallel storage is available, and what throughput can it sustain for our job sizes?
  • What power density and cooling can you support per rack, and is liquid cooling available?
  • How is capacity priced: reserved, on demand, per node-hour or per committed term?
  • Is capacity guaranteed when we need it, or subject to availability?
  • How do we get large data sets in and out, and what are the transfer and egress costs?
  • Do you support our job scheduler and applications, and who handles cluster administration?

How it differs from GPU as a service

GPU as a service is a way of renting GPU capacity from a provider instead of buying the hardware. HPC is the broader discipline of running computationally intensive workloads on high-performance systems, on CPUs, GPUs or both, regardless of who owns the hardware. A rented GPU cluster with the interconnect, storage and scheduling to run parallel jobs is a typical HPC system, and a single powerful rented GPU instance can also serve an HPC workload; what GPUaaS describes is the rental model, not the workload. Many organizations combine the two, owning a base HPC cluster and renting GPU or cloud capacity for peaks.

Frequently Asked Questions

Is HPC the same as supercomputing?
Closely related. Supercomputers are the largest HPC systems, usually run by national labs and research centers. HPC is the broader practice of solving big computing problems with high-performance systems, and it includes the much smaller clusters many companies and universities run.
Can HPC run in the public cloud?
Yes. Major cloud providers offer instances, networking and storage designed for HPC, and many organizations use them for bursts of demand or to avoid buying hardware. Steady, heavily used workloads can cost more in the cloud than on owned or hosted hardware, so compare on your actual usage pattern.
Is HPC the same as AI infrastructure?
They overlap. Training large AI models uses many of the same building blocks as HPC, such as GPU clusters, fast interconnects and parallel storage. Traditional HPC also covers simulation and modeling work that may run mainly on CPUs. The facility and networking requirements are often similar.
Who uses HPC outside of research?
Engineering firms running simulations, pharmaceutical and life sciences companies, financial firms running risk models, energy companies doing seismic analysis, media companies rendering video, and manufacturers using digital design and testing are common examples.

You Don’t Need Another Sales Call. You Need an Answer.

30 minutes. No pitch. Just an honest conversation about where you are, what you need, and whether working together makes sense.

We use your details to set up and prepare for the call, and send the newsletter only if you ask for it. Privacy policy.