Xeon Platinum 8380 AVX-512 Workloads

Benchmarks for a future article. 2 x Intel Xeon Platinum 8380 testing with a Intel M50CYP2SB2U (SE5C6200.86B.0022.D08.2103221623 BIOS) and ASPEED on Ubuntu 22.10 via the Phoronix Test Suite.

HTML result view exported from: https://openbenchmarking.org/result/2308099-NE-XEONPLATI49&grw&sro.

SPECFEM3D

Model: Homogeneous Halfspace

miniBUDE

Implementation: OpenMP - Input Deck: BM1

HeFFTe - Highly Efficient FFT for Exascale

Test: c2c - Backend: FFTW - Precision: double - X Y Z: 128

HeFFTe - Highly Efficient FFT for Exascale

Test: c2c - Backend: FFTW - Precision: double - X Y Z: 256

HeFFTe - Highly Efficient FFT for Exascale

Test: r2c - Backend: FFTW - Precision: float - X Y Z: 128

HeFFTe - Highly Efficient FFT for Exascale

Test: r2c - Backend: FFTW - Precision: float - X Y Z: 256

Palabos

Grid Size: 500

HeFFTe - Highly Efficient FFT for Exascale

Test: c2c - Backend: FFTW - Precision: float - X Y Z: 256

Laghos

Test: Triple Point Problem

Palabos

Grid Size: 100

libxsmm

M N K: 32

libxsmm

M N K: 64

libxsmm

M N K: 256

libxsmm

M N K: 128

miniBUDE

Implementation: OpenMP - Input Deck: BM1

HeFFTe - Highly Efficient FFT for Exascale

Test: r2c - Backend: FFTW - Precision: double - X Y Z: 128

Laghos

Test: Sedov Blast Wave, ube_922_hex.mesh

HeFFTe - Highly Efficient FFT for Exascale

Test: r2c - Backend: FFTW - Precision: double - X Y Z: 256

HeFFTe - Highly Efficient FFT for Exascale

Test: c2c - Backend: FFTW - Precision: float - X Y Z: 128

Palabos

Grid Size: 400

SPECFEM3D

Model: Water-layered Halfspace

miniBUDE

Implementation: OpenMP - Input Deck: BM2

SPECFEM3D

Model: Tomographic Model

SPECFEM3D

Model: Layered Halfspace

SPECFEM3D

Model: Mount St. Helens

miniBUDE

Implementation: OpenMP - Input Deck: BM2

Timed MrBayes Analysis

Primate Phylogeny Analysis

TensorFlow

Device: CPU - Batch Size: 256 - Model: AlexNet

TensorFlow

Device: CPU - Batch Size: 512 - Model: AlexNet

Remhos

Test: Sample Remap Example

TensorFlow

Device: CPU - Batch Size: 256 - Model: GoogLeNet

TensorFlow

Device: CPU - Batch Size: 256 - Model: ResNet-50

TensorFlow

Device: CPU - Batch Size: 512 - Model: GoogLeNet

TensorFlow

Device: CPU - Batch Size: 512 - Model: ResNet-50

CloverLeaf

Lagrangian-Eulerian Hydrodynamics

Neural Magic DeepSparse

Model: NLP Document Classification, oBERT base uncased on IMDB - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: NLP Document Classification, oBERT base uncased on IMDB - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: NLP Text Classification, BERT base uncased SST2, Sparse INT8 - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: NLP Text Classification, BERT base uncased SST2, Sparse INT8 - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: NLP Sentiment Analysis, 80% Pruned Quantized BERT Base Uncased - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: NLP Sentiment Analysis, 80% Pruned Quantized BERT Base Uncased - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: NLP Question Answering, BERT base uncased SQuaD 12layer Pruned90 - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: NLP Question Answering, BERT base uncased SQuaD 12layer Pruned90 - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: ResNet-50, Baseline - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: ResNet-50, Baseline - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: ResNet-50, Sparse INT8 - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: ResNet-50, Sparse INT8 - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: CV Detection, YOLOv5s COCO - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: CV Detection, YOLOv5s COCO - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: BERT-Large, NLP Question Answering - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: BERT-Large, NLP Question Answering - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: CV Classification, ResNet-50 ImageNet - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: CV Classification, ResNet-50 ImageNet - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: CV Detection, YOLOv5s COCO, Sparse INT8 - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: CV Detection, YOLOv5s COCO, Sparse INT8 - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: NLP Text Classification, DistilBERT mnli - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: NLP Text Classification, DistilBERT mnli - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: CV Segmentation, 90% Pruned YOLACT Pruned - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: CV Segmentation, 90% Pruned YOLACT Pruned - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: BERT-Large, NLP Question Answering, Sparse INT8 - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: BERT-Large, NLP Question Answering, Sparse INT8 - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: NLP Text Classification, BERT base uncased SST2 - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: NLP Text Classification, BERT base uncased SST2 - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: NLP Token Classification, BERT base uncased conll2003 - Scenario: Asynchronous Multi-Stream

Neural Magic DeepSparse

Model: NLP Token Classification, BERT base uncased conll2003 - Scenario: Asynchronous Multi-Stream

ONNX Runtime

Model: GPT-2 - Device: CPU - Executor: Standard

ONNX Runtime

Model: yolov4 - Device: CPU - Executor: Standard

ONNX Runtime

Model: bertsquad-12 - Device: CPU - Executor: Standard

ONNX Runtime

Model: CaffeNet 12-int8 - Device: CPU - Executor: Standard

ONNX Runtime

Model: fcn-resnet101-11 - Device: CPU - Executor: Standard

ONNX Runtime

Model: ArcFace ResNet-100 - Device: CPU - Executor: Standard

ONNX Runtime

Model: ResNet50 v1-12-int8 - Device: CPU - Executor: Standard

ONNX Runtime

Model: super-resolution-10 - Device: CPU - Executor: Standard

NCNN

Target: CPU - Model: mobilenet

NCNN

Target: CPU-v2-v2 - Model: mobilenet-v2

NCNN

Target: CPU-v3-v3 - Model: mobilenet-v3

NCNN

Target: CPU - Model: shufflenet-v2

NCNN

Target: CPU - Model: mnasnet

NCNN

Target: CPU - Model: efficientnet-b0

NCNN

Target: CPU - Model: blazeface

NCNN

Target: CPU - Model: googlenet

NCNN

Target: CPU - Model: vgg16

NCNN

Target: CPU - Model: resnet18

NCNN

Target: CPU - Model: alexnet

NCNN

Target: CPU - Model: resnet50

NCNN

Target: CPU - Model: yolov4-tiny

NCNN

Target: CPU - Model: squeezenet_ssd

NCNN

Target: CPU - Model: regnety_400m

NCNN

Target: CPU - Model: vision_transformer

NCNN

Target: CPU - Model: FastestDet

GROMACS

Implementation: MPI CPU - Input: water_GMX50_bare

oneDNN

Harness: IP Shapes 3D - Data Type: bf16bf16bf16 - Engine: CPU

oneDNN

Harness: Convolution Batch Shapes Auto - Data Type: bf16bf16bf16 - Engine: CPU

oneDNN

Harness: Deconvolution Batch shapes_1d - Data Type: bf16bf16bf16 - Engine: CPU

oneDNN

Harness: Deconvolution Batch shapes_3d - Data Type: bf16bf16bf16 - Engine: CPU

oneDNN

Harness: Recurrent Neural Network Training - Data Type: bf16bf16bf16 - Engine: CPU

oneDNN

Harness: Recurrent Neural Network Inference - Data Type: bf16bf16bf16 - Engine: CPU

OpenVINO

Model: Face Detection FP16 - Device: CPU

OpenVINO

Model: Face Detection FP16 - Device: CPU

OpenVINO

Model: Person Detection FP16 - Device: CPU

OpenVINO

Model: Person Detection FP16 - Device: CPU

OpenVINO

Model: Person Detection FP32 - Device: CPU

OpenVINO

Model: Person Detection FP32 - Device: CPU

OpenVINO

Model: Vehicle Detection FP16 - Device: CPU

OpenVINO

Model: Vehicle Detection FP16 - Device: CPU

OpenVINO

Model: Face Detection FP16-INT8 - Device: CPU

OpenVINO

Model: Face Detection FP16-INT8 - Device: CPU

OpenVINO

Model: Vehicle Detection FP16-INT8 - Device: CPU

OpenVINO

Model: Vehicle Detection FP16-INT8 - Device: CPU

OpenVINO

Model: Weld Porosity Detection FP16 - Device: CPU

OpenVINO

Model: Weld Porosity Detection FP16 - Device: CPU

OpenVINO

Model: Machine Translation EN To DE FP16 - Device: CPU

OpenVINO

Model: Machine Translation EN To DE FP16 - Device: CPU

OpenVINO

Model: Weld Porosity Detection FP16-INT8 - Device: CPU

OpenVINO

Model: Weld Porosity Detection FP16-INT8 - Device: CPU

OpenVINO

Model: Person Vehicle Bike Detection FP16 - Device: CPU

OpenVINO

Model: Person Vehicle Bike Detection FP16 - Device: CPU

OpenVINO

Model: Age Gender Recognition Retail 0013 FP16 - Device: CPU

OpenVINO

Model: Age Gender Recognition Retail 0013 FP16 - Device: CPU

OpenVINO

Model: Age Gender Recognition Retail 0013 FP16-INT8 - Device: CPU

OpenVINO

Model: Age Gender Recognition Retail 0013 FP16-INT8 - Device: CPU

QMCPACK

Input: Li2_STO_ae

QMCPACK

Input: simple-H2O

QMCPACK

Input: FeCO6_b3lyp_gms

QMCPACK

Input: FeCO6_b3lyp_gms

Xcompact3d Incompact3d

Input: input.i3d 193 Cells Per Direction

Cpuminer-Opt

Algorithm: Magi

Cpuminer-Opt

Algorithm: x25x

Cpuminer-Opt

Algorithm: scrypt

Cpuminer-Opt

Algorithm: Deepcoin

Cpuminer-Opt

Algorithm: Blake-2 S

Cpuminer-Opt

Algorithm: Garlicoin

Cpuminer-Opt

Algorithm: Skeincoin

Cpuminer-Opt

Algorithm: Myriad-Groestl

Cpuminer-Opt

Algorithm: LBC, LBRY Credits

Cpuminer-Opt

Algorithm: Quad SHA-256, Pyrite

Cpuminer-Opt

Algorithm: Triple SHA-256, Onecoin

VP9 libvpx Encoding

Speed: Speed 5 - Input: Bosphorus 4K

dav1d

Video Input: Chimera 1080p

dav1d

Video Input: Summer Nature 4K

SVT-AV1

Encoder Mode: Preset 8 - Input: Bosphorus 4K

SVT-AV1

Encoder Mode: Preset 12 - Input: Bosphorus 4K

SVT-AV1

Encoder Mode: Preset 13 - Input: Bosphorus 4K

SVT-HEVC

Tuning: 1 - Input: Bosphorus 4K

SVT-HEVC

Tuning: 7 - Input: Bosphorus 4K

SVT-HEVC

Tuning: 10 - Input: Bosphorus 4K

Blender

Blend File: BMW27 - Compute: CPU-Only

Blender

Blend File: Fishy Cat - Compute: CPU-Only

VVenC

Video Input: Bosphorus 4K - Video Preset: Fast

VVenC

Video Input: Bosphorus 4K - Video Preset: Faster

Embree

Binary: Pathtracer ISPC - Model: Crown

Embree

Binary: Pathtracer ISPC - Model: Asian Dragon

Intel Open Image Denoise

Run: RT.ldr_alb_nrm.3840x2160 - Device: CPU-Only

Intel Open Image Denoise

Run: RTLightmap.hdr.4096x4096 - Device: CPU-Only

OpenVKL

Benchmark: vklBenchmark ISPC

OSPRay

Benchmark: particle_volume/ao/real_time

OSPRay

Benchmark: particle_volume/scivis/real_time

OSPRay

Benchmark: particle_volume/pathtracer/real_time

OSPRay

Benchmark: gravity_spheres_volume/dim_512/ao/real_time

OSPRay

Benchmark: gravity_spheres_volume/dim_512/scivis/real_time

simdjson

Throughput Test: Kostya

simdjson

Throughput Test: TopTweet

simdjson

Throughput Test: LargeRandom

simdjson

Throughput Test: PartialTweets

simdjson

Throughput Test: DistinctUserID

ONNX Runtime

Model: GPT-2 - Device: CPU - Executor: Standard

ONNX Runtime

Model: yolov4 - Device: CPU - Executor: Standard

ONNX Runtime

Model: bertsquad-12 - Device: CPU - Executor: Standard

ONNX Runtime

Model: CaffeNet 12-int8 - Device: CPU - Executor: Standard

ONNX Runtime

Model: fcn-resnet101-11 - Device: CPU - Executor: Standard

ONNX Runtime

Model: ArcFace ResNet-100 - Device: CPU - Executor: Standard

ONNX Runtime

Model: ResNet50 v1-12-int8 - Device: CPU - Executor: Standard

ONNX Runtime

Model: super-resolution-10 - Device: CPU - Executor: Standard

Phoronix Test Suite v10.8.5