OMS Reference
Understand the core concepts of high-performance computing, artificial intelligence, data centers and energy, then clearly tell HPC and AI infrastructure apart from ASIC mining hardware.
HPC and Artificial Intelligence
HPC
High-Performance Computing: computing performed on systems designed to process scientific, industrial or engineering workloads quickly.
Artificial Intelligence (AI)
A set of methods that enable a system to perform tasks such as classification, content generation or decision-making.
Machine Learning
Learning automatically from data, without explicitly programming every decision rule.
Deep Learning
A subfield of machine learning that uses neural networks with many layers.
Training
The phase in which a model learns from a dataset; it often requires a lot of compute and memory.
Inference
Using an already trained model to produce a prediction or a response.
Model
A mathematical structure trained to perform a specific task.
Parameter
An internal value learned by a model. The number of parameters affects, among other things, the memory and resources required.
Hardware and Acceleration
CPU
A general-purpose processor suited to a wide range of sequential computations and control tasks.
GPU
A massively parallel processor used for graphics rendering, AI and some HPC workloads.
ASIC
An integrated circuit designed for one specific task. A mining ASIC runs a given algorithm efficiently but cannot replace a general-purpose GPU.
FPGA
A reconfigurable circuit whose hardware logic can be adapted to a particular workload.
Accelerator
A specialized component that offloads specific computations from the CPU, such as a GPU or an ASIC.
VRAM
A GPU’s onboard memory. Its capacity and bandwidth limit the size of the models and data batches it can process.
Memory Bandwidth
The amount of data memory can transfer per second. It can become the limiting factor before compute power does.
Interconnect
The link between accelerators, servers or racks. Its latency and throughput have a major impact on distributed workloads.
Infrastructure and Data Centers
Server
A computer system designed to provide compute, storage or network resources on a continuous basis.
Compute Node
An individual machine that is part of an HPC or AI cluster.
Cluster
A group of coordinated nodes that handles larger workloads or provides higher availability.
Rack
A standardized enclosure that houses servers, networking equipment and power distribution.
PDU
Power Distribution Unit: equipment that distributes electrical power to the devices in a rack or an installation.
Redundancy
Duplicating critical components, such as power or networking, to reduce outages.
Latency
The time needed to transmit information or get a response.
Network Throughput
The amount of data transferred between systems per unit of time.
Energy and Cooling
Power Draw (W)
The instantaneous demand of a piece of equipment. It determines how circuits, circuit protection, cables and PDUs are sized.
Energy Consumption (kWh)
The energy used over a period of time. It is used to calculate the actual electricity cost.
PUE
Power Usage Effectiveness: the ratio of a site’s total energy use to the energy used by its IT equipment. A value close to 1 means lower overhead losses.
Air Cooling
Removing heat with airflow, fans and an air-conditioning or exhaust system.
Liquid Cooling
Transferring heat through a fluid circulating in cold plates or a dedicated loop.
Immersion Cooling
Cooling equipment by submerging it in a compatible dielectric fluid.
Heat Density
The amount of heat concentrated in a given area or volume. It influences the choice of cooling.
Heat Recovery
Reusing the heat produced to warm a building, a fluid or a process.
Operations and Performance
Uptime
The share of time a system remains available and operational.
Monitoring
Tracking performance, temperatures, power consumption, errors and availability status.
Telemetry
Measurements sent by equipment to a monitoring or alerting system.
Workload
The set of computations and data assigned to an infrastructure.
Load Balancing
Spreading a workload across multiple resources to avoid bottlenecks.
Orchestrator
A tool that schedules, deploys and supervises tasks or services across multiple machines.
SLA
Service Level Agreement: a measurable commitment to availability or service level.
TCO
Total Cost of Ownership: the full cost, including purchase, energy, cooling, maintenance, networking and operations.
Mining, HPC and AI: How to Compare
Proof of Work (PoW)
A consensus mechanism in which machines perform verifiable computations to secure a blockchain network.
Hashrate
The number of hash calculations performed per second by a miner or a network.
Mining Profitability
The difference between estimated mining revenue and costs, including electricity, hosting and hardware purchases.
Performance per Watt
Performance measured relative to the power consumed. The relevant unit depends on the workload and the hardware.
Specialized Hardware
Equipment optimized for a defined task. Higher efficiency generally comes at the cost of versatility.
Hardware Repurposing
Putting equipment to a different use. It is limited by the architecture: a mining ASIC generally cannot be converted into a GPU server.
CapEx
Upfront capital expenditures, such as buying machines, racks, transformers and cooling systems.
OpEx
Recurring operating expenses, including electricity, connectivity, maintenance, staff and hosting.
Connecting HPC, AI and ASIC Mining
Use OMS guides and tools to dig deeper into hardware, operations and profitability.