Intel Core i9-13900K testing with a ASUS TUF GAMING Z790-PRO WIFI (1630 BIOS) and ASUS NVIDIA GeForce RTX 4070 Ti SUPER 16GB on EndeavourOS rolling via the Phoronix Test Suite.
NVIDIA RTX 4070 SUPER Processor: Intel Core i9-13900K @ 5.50GHz (24 Cores / 32 Threads), Motherboard: ASUS TUF GAMING Z790-PRO WIFI (1401 BIOS), Chipset: Intel Device 7a27, Memory: 32GB, Disk: 4001GB Seagate ZP4000GP304001, Graphics: ASUS NVIDIA GeForce RTX 4070 SUPER 12GB, Audio: Realtek ALC1220, Monitor: ARZOPA, Network: Intel I226-V + Intel Device 7a70
OS: EndeavourOS rolling, Kernel: 6.7.1-arch1-1 (x86_64), Desktop: KDE Plasma 5.27.10, Display Server: X Server 1.21.1.11, Display Driver: NVIDIA 550.40.07, OpenGL: 4.6.0, OpenCL: OpenCL 3.0 CUDA 12.4.74, Compiler: GCC 13.2.1 20230801, File-System: ext4, Screen Resolution: 1920x1080
Kernel Notes: Transparent Huge Pages: alwaysCompiler Notes: --disable-libssp --disable-libstdcxx-pch --disable-werror --enable-__cxa_atexit --enable-bootstrap --enable-cet=auto --enable-checking=release --enable-clocale=gnu --enable-default-pie --enable-default-ssp --enable-gnu-indirect-function --enable-gnu-unique-object --enable-languages=ada,c,c++,d,fortran,go,lto,objc,obj-c++ --enable-libstdcxx-backtrace --enable-link-serialization=1 --enable-lto --enable-multilib --enable-plugin --enable-shared --enable-threads=posix --mandir=/usr/share/man --with-build-config=bootstrap-lto --with-linker-hash-style=gnuProcessor Notes: Scaling Governor: intel_pstate powersave (EPP: balance_performance) - CPU Microcode: 0x11dGraphics Notes: BAR1 / Visible vRAM Size: 16384 MiB - vBIOS Version: 95.04.69.00.c1Security Notes: gather_data_sampling: Not affected + itlb_multihit: Not affected + l1tf: Not affected + mds: Not affected + meltdown: Not affected + mmio_stale_data: Not affected + retbleed: Not affected + spec_rstack_overflow: Not affected + spec_store_bypass: Mitigation of SSB disabled via prctl + spectre_v1: Mitigation of usercopy/swapgs barriers and __user pointer sanitization + spectre_v2: Mitigation of Enhanced / Automatic IBRS IBPB: conditional RSB filling PBRSB-eIBRS: SW sequence + srbds: Not affected + tsx_async_abort: Not affected
NVIDIA RTX 4070 Processor: Intel Core i9-13900K @ 5.50GHz (24 Cores / 32 Threads), Motherboard: ASUS TUF GAMING Z790-PRO WIFI (1401 BIOS), Chipset: Intel Device 7a27, Memory: 32GB, Disk: 4001GB Seagate ZP4000GP304001, Graphics: MSI NVIDIA GeForce RTX 4070 12GB , Audio: Realtek ALC1220, Monitor: ARZOPA, Network: Intel I226-V + Intel Device 7a70
OS: EndeavourOS rolling, Kernel: 6.7.1-arch1-1 (x86_64), Desktop: KDE Plasma 5.27.10, Display Server: X Server 1.21.1.11, Display Driver: NVIDIA 550.40.07, OpenGL: 4.6.0, OpenCL: OpenCL 3.0 CUDA 12.4.74, Compiler: GCC 13.2.1 20230801 + CUDA 12.3, File-System: ext4, Screen Resolution: 1920x1080
Kernel Notes: Transparent Huge Pages: alwaysEnvironment Notes: NVCC_PREPEND_FLAGS="-ccbin /opt/cuda/bin"Compiler Notes: --disable-libssp --disable-libstdcxx-pch --disable-werror --enable-__cxa_atexit --enable-bootstrap --enable-cet=auto --enable-checking=release --enable-clocale=gnu --enable-default-pie --enable-default-ssp --enable-gnu-indirect-function --enable-gnu-unique-object --enable-languages=ada,c,c++,d,fortran,go,lto,objc,obj-c++ --enable-libstdcxx-backtrace --enable-link-serialization=1 --enable-lto --enable-multilib --enable-plugin --enable-shared --enable-threads=posix --mandir=/usr/share/man --with-build-config=bootstrap-lto --with-linker-hash-style=gnuProcessor Notes: Scaling Governor: intel_pstate powersave (EPP: balance_performance) - CPU Microcode: 0x11dGraphics Notes: BAR1 / Visible vRAM Size: 16384 MiB - vBIOS Version: 95.04.3e.40.2aPython Notes: Python 3.11.6Security Notes: gather_data_sampling: Not affected + itlb_multihit: Not affected + l1tf: Not affected + mds: Not affected + meltdown: Not affected + mmio_stale_data: Not affected + retbleed: Not affected + spec_rstack_overflow: Not affected + spec_store_bypass: Mitigation of SSB disabled via prctl + spectre_v1: Mitigation of usercopy/swapgs barriers and __user pointer sanitization + spectre_v2: Mitigation of Enhanced / Automatic IBRS IBPB: conditional RSB filling PBRSB-eIBRS: SW sequence + srbds: Not affected + tsx_async_abort: Not affected
NVIDIA RTX 4070 TI Changed Graphics to NVIDIA GeForce RTX 4070 Ti 12GB .
Graphics Change: BAR1 / Visible vRAM Size: 16384 MiB - vBIOS Version: 95.04.31.00.36
NVIDIA RTX 3090 Processor: Intel Core i9-13900K @ 5.50GHz (24 Cores / 32 Threads), Motherboard: ASUS TUF GAMING Z790-PRO WIFI (1401 BIOS), Chipset: Intel Device 7a27, Memory: 32GB, Disk: 4001GB Seagate ZP4000GP304001, Graphics: NVIDIA GeForce RTX 3090 24GB , Audio: Realtek ALC1220, Monitor: PI-KVM Video , Network: Intel I226-V + Intel Device 7a70
OS: EndeavourOS rolling, Kernel: 6.7.4-arch1-1 (x86_64), Desktop: KDE Plasma 5.27.10, Display Server: X Server 1.21.1.11, Display Driver: NVIDIA 550.40.07, OpenGL: 4.6.0, OpenCL: OpenCL 3.0 CUDA 12.4.74, Compiler: GCC 13.2.1 20230801 + CUDA 12.3, File-System: ext4, Screen Resolution: 1920x1080
Kernel Notes: Transparent Huge Pages: alwaysEnvironment Notes: NVCC_PREPEND_FLAGS="-ccbin /opt/cuda/bin"Compiler Notes: --disable-libssp --disable-libstdcxx-pch --disable-werror --enable-__cxa_atexit --enable-bootstrap --enable-cet=auto --enable-checking=release --enable-clocale=gnu --enable-default-pie --enable-default-ssp --enable-gnu-indirect-function --enable-gnu-unique-object --enable-languages=ada,c,c++,d,fortran,go,lto,m2,objc,obj-c++ --enable-libstdcxx-backtrace --enable-link-serialization=1 --enable-lto --enable-multilib --enable-plugin --enable-shared --enable-threads=posix --mandir=/usr/share/man --with-build-config=bootstrap-lto --with-linker-hash-style=gnuProcessor Notes: Scaling Governor: intel_pstate powersave (EPP: balance_performance) - CPU Microcode: 0x11dGraphics Notes: BAR1 / Visible vRAM Size: 256 MiB - vBIOS Version: 94.02.26.08.baPython Notes: Python 3.11.6Security Notes: gather_data_sampling: Not affected + itlb_multihit: Not affected + l1tf: Not affected + mds: Not affected + meltdown: Not affected + mmio_stale_data: Not affected + retbleed: Not affected + spec_rstack_overflow: Not affected + spec_store_bypass: Mitigation of SSB disabled via prctl + spectre_v1: Mitigation of usercopy/swapgs barriers and __user pointer sanitization + spectre_v2: Mitigation of Enhanced / Automatic IBRS IBPB: conditional RSB filling PBRSB-eIBRS: SW sequence + srbds: Not affected + tsx_async_abort: Not affected
NVIDIA RTX 4070 TI SUPER Processor: Intel Core i9-13900K @ 5.50GHz (24 Cores / 32 Threads), Motherboard: ASUS TUF GAMING Z790-PRO WIFI (1630 BIOS) , Chipset: Intel Raptor Lake-S PCH , Memory: 32GB, Disk: 4001GB Seagate ZP4000GP304001 + 0GB CD-ROM Drive , Graphics: ASUS NVIDIA GeForce RTX 4070 Ti SUPER 16GB , Audio: Realtek ALC1220, Monitor: PI-KVM Video, Network: Intel I226-V + Intel Raptor Lake-S PCH CNVi WiFi
OS: EndeavourOS rolling, Kernel: 6.7.4-arch1-1 (x86_64), Desktop: KDE Plasma 5.27.10, Display Server: X Server 1.21.1.11, Display Driver: NVIDIA 550.40.07, OpenGL: 4.6.0, OpenCL: OpenCL 2.1 AMD-APP (3602.0) + OpenCL 3.0 CUDA 12.4.74, Compiler: GCC 13.2.1 20230801 + CUDA 12.3, File-System: ext4, Screen Resolution: 1920x1080
Kernel Notes: Transparent Huge Pages: alwaysEnvironment Notes: NVCC_PREPEND_FLAGS="-ccbin /opt/cuda/bin"Compiler Notes: --disable-libssp --disable-libstdcxx-pch --disable-werror --enable-__cxa_atexit --enable-bootstrap --enable-cet=auto --enable-checking=release --enable-clocale=gnu --enable-default-pie --enable-default-ssp --enable-gnu-indirect-function --enable-gnu-unique-object --enable-languages=ada,c,c++,d,fortran,go,lto,m2,objc,obj-c++ --enable-libstdcxx-backtrace --enable-link-serialization=1 --enable-lto --enable-multilib --enable-plugin --enable-shared --enable-threads=posix --mandir=/usr/share/man --with-build-config=bootstrap-lto --with-linker-hash-style=gnuProcessor Notes: Scaling Governor: intel_pstate powersave (EPP: balance_performance) - CPU Microcode: 0x11fGraphics Notes: BAR1 / Visible vRAM Size: 256 MiB - vBIOS Version: 95.03.45.00.c5Python Notes: Python 3.11.7Security Notes: gather_data_sampling: Not affected + itlb_multihit: Not affected + l1tf: Not affected + mds: Not affected + meltdown: Not affected + mmio_stale_data: Not affected + retbleed: Not affected + spec_rstack_overflow: Not affected + spec_store_bypass: Mitigation of SSB disabled via prctl + spectre_v1: Mitigation of usercopy/swapgs barriers and __user pointer sanitization + spectre_v2: Mitigation of Enhanced / Automatic IBRS IBPB: conditional RSB filling PBRSB-eIBRS: SW sequence + srbds: Not affected + tsx_async_abort: Not affected
TensorFlow This is a benchmark of the TensorFlow deep learning framework using the TensorFlow reference benchmarks (tensorflow/benchmarks with tf_cnn_benchmarks.py). Note with the Phoronix Test Suite there is also pts/tensorflow-lite for benchmarking the TensorFlow Lite binaries if desired for complementary metrics. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org images/sec, More Is Better TensorFlow 2.12 Device: GPU - Batch Size: 1 - Model: AlexNet NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 4 8 12 16 20 SE +/- 0.22, N = 2 SE +/- 0.16, N = 3 SE +/- 0.06, N = 2 SE +/- 0.20, N = 15 SE +/- 0.13, N = 15 13.92 14.04 14.79 14.45 12.26
OpenBenchmarking.org images/sec, More Is Better TensorFlow 2.12 Device: GPU - Batch Size: 16 - Model: VGG-16 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 0.3375 0.675 1.0125 1.35 1.6875 SE +/- 0.00, N = 2 SE +/- 0.01, N = 2 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 1.48 1.50 1.49 1.49 1.45
OpenBenchmarking.org images/sec, More Is Better TensorFlow 2.12 Device: GPU - Batch Size: 32 - Model: VGG-16 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 0.3375 0.675 1.0125 1.35 1.6875 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.00, N = 2 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 1.50 1.50 1.50 1.50 1.46
OpenBenchmarking.org images/sec, More Is Better TensorFlow 2.12 Device: GPU - Batch Size: 64 - Model: VGG-16 NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 0.3398 0.6796 1.0194 1.3592 1.699 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 1.50 1.50 1.51 1.46
Device: GPU - Batch Size: 64 - Model: VGG-16
NVIDIA RTX 4070 SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: UnboundLocalError: cannot access local variable 'decorators' where it is not associated with a value
OpenBenchmarking.org images/sec, More Is Better TensorFlow 2.12 Device: GPU - Batch Size: 16 - Model: AlexNet NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 7 14 21 28 35 SE +/- 0.17, N = 3 SE +/- 0.08, N = 3 SE +/- 0.07, N = 3 SE +/- 0.07, N = 3 31.59 31.45 31.70 31.98 31.10
OpenBenchmarking.org images/sec, More Is Better TensorFlow 2.12 Device: GPU - Batch Size: 32 - Model: AlexNet NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 8 16 24 32 40 SE +/- 0.15, N = 2 SE +/- 0.18, N = 3 SE +/- 0.04, N = 3 SE +/- 0.05, N = 3 SE +/- 0.19, N = 3 33.40 33.32 33.29 33.53 32.88
ProjectPhysX OpenCL-Benchmark ProjectPhysX OpenCL-Benchmark provides various OpenCL compute and memory bandwidth micro-benchmarks Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org GB/s, More Is Better ProjectPhysX OpenCL-Benchmark 1.2 Operation: Memory Bandwidth Coalesced Write NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 200 400 600 800 1000 SE +/- 0.14, N = 3 SE +/- 0.16, N = 3 SE +/- 0.11, N = 3 SE +/- 0.06, N = 3 SE +/- 0.57, N = 3 455.01 459.43 457.17 887.31 608.94 1. (CXX) g++ options: -std=c++17 -pthread -lOpenCL
TensorFlow This is a benchmark of the TensorFlow deep learning framework using the TensorFlow reference benchmarks (tensorflow/benchmarks with tf_cnn_benchmarks.py). Note with the Phoronix Test Suite there is also pts/tensorflow-lite for benchmarking the TensorFlow Lite binaries if desired for complementary metrics. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org images/sec, More Is Better TensorFlow 2.12 Device: GPU - Batch Size: 1 - Model: VGG-16 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 0.3105 0.621 0.9315 1.242 1.5525 SE +/- 0.01, N = 2 SE +/- 0.01, N = 3 SE +/- 0.01, N = 3 SE +/- 0.00, N = 3 1.35 1.36 1.38 1.38 1.32
ProjectPhysX OpenCL-Benchmark ProjectPhysX OpenCL-Benchmark provides various OpenCL compute and memory bandwidth micro-benchmarks Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org TIOPs/s, More Is Better ProjectPhysX OpenCL-Benchmark 1.2 Operation: INT8 Compute NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 4 8 12 16 20 SE +/- 0.05, N = 3 SE +/- 0.02, N = 3 SE +/- 0.03, N = 3 SE +/- 0.07, N = 3 SE +/- 0.00, N = 3 14.31 12.12 15.73 13.73 17.62 1. (CXX) g++ options: -std=c++17 -pthread -lOpenCL
OpenBenchmarking.org GB/s, More Is Better ProjectPhysX OpenCL-Benchmark 1.2 Operation: Memory Bandwidth Coalesced Read NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 200 400 600 800 1000 SE +/- 0.01, N = 3 SE +/- 0.03, N = 3 SE +/- 0.01, N = 3 SE +/- 0.07, N = 3 SE +/- 0.06, N = 3 464.86 465.18 465.07 864.11 619.03 1. (CXX) g++ options: -std=c++17 -pthread -lOpenCL
OpenBenchmarking.org TIOPs/s, More Is Better ProjectPhysX OpenCL-Benchmark 1.2 Operation: INT16 Compute NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 5 10 15 20 25 SE +/- 0.00, N = 3 SE +/- 0.02, N = 3 SE +/- 0.02, N = 3 SE +/- 0.00, N = 3 SE +/- 0.02, N = 3 17.17 14.28 18.28 17.00 20.50 1. (CXX) g++ options: -std=c++17 -pthread -lOpenCL
OpenBenchmarking.org TIOPs/s, More Is Better ProjectPhysX OpenCL-Benchmark 1.2 Operation: INT32 Compute NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 6 12 18 24 30 SE +/- 0.00, N = 3 SE +/- 0.02, N = 3 SE +/- 0.04, N = 3 SE +/- 0.06, N = 3 SE +/- 0.01, N = 3 19.89 16.38 21.05 20.03 23.66 1. (CXX) g++ options: -std=c++17 -pthread -lOpenCL
OpenBenchmarking.org TIOPs/s, More Is Better ProjectPhysX OpenCL-Benchmark 1.2 Operation: INT64 Compute NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 0.9945 1.989 2.9835 3.978 4.9725 SE +/- 0.015, N = 3 SE +/- 0.004, N = 3 SE +/- 0.016, N = 3 SE +/- 0.003, N = 3 SE +/- 0.009, N = 3 4.214 3.443 4.420 3.135 4.414 1. (CXX) g++ options: -std=c++17 -pthread -lOpenCL
TensorFlow This is a benchmark of the TensorFlow deep learning framework using the TensorFlow reference benchmarks (tensorflow/benchmarks with tf_cnn_benchmarks.py). Note with the Phoronix Test Suite there is also pts/tensorflow-lite for benchmarking the TensorFlow Lite binaries if desired for complementary metrics. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org images/sec, More Is Better TensorFlow 2.12 Device: GPU - Batch Size: 256 - Model: VGG-16 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 0.3398 0.6796 1.0194 1.3592 1.699 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 1.50 1.51 1.47
Device: GPU - Batch Size: 256 - Model: VGG-16
NVIDIA RTX 4070 SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status.
NVIDIA RTX 4070: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: AttributeError: 'collections.OrderedDict' object has no attribute 'empty'
Device: GPU - Batch Size: 512 - Model: VGG-16
NVIDIA RTX 4070 SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: Fatal Python error: Segmentation fault
NVIDIA RTX 4070: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: AttributeError: 'function' object has no attribute 'empty'
NVIDIA RTX 4070 TI: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: Fatal Python error: Segmentation fault
NVIDIA RTX 3090: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: Fatal Python error: Segmentation fault
NVIDIA RTX 4070 TI SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: Fatal Python error: Segmentation fault
OpenBenchmarking.org images/sec, More Is Better TensorFlow 2.12 Device: GPU - Batch Size: 64 - Model: AlexNet NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 8 16 24 32 40 SE +/- 0.14, N = 3 SE +/- 0.06, N = 3 SE +/- 0.08, N = 3 SE +/- 0.06, N = 3 33.97 33.93 34.06 33.93 33.55
GpuOwl GpuOwl is a Mersenne primality tester leveraging OpenCL for cross-vendor GPU acceleration. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org Iterations / Second, More Is Better GpuOwl 7.2.1 Exponent: 77936867 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 160 320 480 640 800 SE +/- 0.00, N = 3 SE +/- 0.09, N = 3 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 646.41 530.32 676.59 645.99 761.61
OpenBenchmarking.org Iterations / Second, More Is Better GpuOwl 7.2.1 Exponent: 332220523 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 40 80 120 160 200 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.01, N = 3 137.44 112.61 145.84 137.32 163.41
TensorFlow This is a benchmark of the TensorFlow deep learning framework using the TensorFlow reference benchmarks (tensorflow/benchmarks with tf_cnn_benchmarks.py). Note with the Phoronix Test Suite there is also pts/tensorflow-lite for benchmarking the TensorFlow Lite binaries if desired for complementary metrics. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org images/sec, More Is Better TensorFlow 2.12 Device: GPU - Batch Size: 1 - Model: GoogLeNet NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 3 6 9 12 15 SE +/- 0.17, N = 2 SE +/- 0.10, N = 3 SE +/- 0.30, N = 2 SE +/- 0.07, N = 3 SE +/- 0.05, N = 3 12.62 12.78 12.79 12.82 12.24
GpuOwl GpuOwl is a Mersenne primality tester leveraging OpenCL for cross-vendor GPU acceleration. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org Iterations / Second, More Is Better GpuOwl 7.2.1 Exponent: 57885161 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 200 400 600 800 1000 SE +/- 1.26, N = 3 SE +/- 0.00, N = 3 SE +/- 2.53, N = 3 SE +/- 2.01, N = 3 SE +/- 0.35, N = 3 869.07 714.80 919.13 866.31 1025.99
TensorFlow This is a benchmark of the TensorFlow deep learning framework using the TensorFlow reference benchmarks (tensorflow/benchmarks with tf_cnn_benchmarks.py). Note with the Phoronix Test Suite there is also pts/tensorflow-lite for benchmarking the TensorFlow Lite binaries if desired for complementary metrics. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org images/sec, More Is Better TensorFlow 2.12 Device: GPU - Batch Size: 1 - Model: ResNet-50 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 0.9788 1.9576 2.9364 3.9152 4.894 SE +/- 0.01, N = 3 SE +/- 0.02, N = 2 SE +/- 0.03, N = 3 SE +/- 0.02, N = 3 4.35 4.34 4.32 4.35 4.14
OpenBenchmarking.org images/sec, More Is Better TensorFlow 2.12 Device: GPU - Batch Size: 256 - Model: AlexNet NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 8 16 24 32 40 SE +/- 0.01, N = 3 SE +/- 0.07, N = 2 SE +/- 0.07, N = 3 SE +/- 0.05, N = 3 34.16 34.61 34.46 33.95
Device: GPU - Batch Size: 256 - Model: AlexNet
NVIDIA RTX 4070: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: UnboundLocalError: cannot access local variable 'kind' where it is not associated with a value
OpenBenchmarking.org images/sec, More Is Better TensorFlow 2.12 Device: GPU - Batch Size: 512 - Model: AlexNet NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 8 16 24 32 40 SE +/- 0.02, N = 2 SE +/- 0.03, N = 3 SE +/- 0.09, N = 2 SE +/- 0.01, N = 3 SE +/- 0.01, N = 3 35.10 35.21 35.44 35.58 35.02
OpenBenchmarking.org images/sec, More Is Better TensorFlow 2.12 Device: GPU - Batch Size: 16 - Model: GoogLeNet NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 4 8 12 16 20 SE +/- 0.03, N = 3 SE +/- 0.03, N = 3 SE +/- 0.07, N = 3 SE +/- 0.05, N = 3 SE +/- 0.02, N = 3 15.67 15.66 15.69 15.68 15.29
OpenBenchmarking.org images/sec, More Is Better TensorFlow 2.12 Device: GPU - Batch Size: 16 - Model: ResNet-50 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 1.2353 2.4706 3.7059 4.9412 6.1765 SE +/- 0.00, N = 2 SE +/- 0.02, N = 3 SE +/- 0.01, N = 3 SE +/- 0.00, N = 3 SE +/- 0.02, N = 3 5.46 5.49 5.46 5.49 5.32
OpenBenchmarking.org images/sec, More Is Better TensorFlow 2.12 Device: GPU - Batch Size: 32 - Model: GoogLeNet NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 4 8 12 16 20 SE +/- 0.01, N = 2 SE +/- 0.03, N = 3 SE +/- 0.03, N = 3 SE +/- 0.06, N = 3 SE +/- 0.06, N = 3 15.61 15.63 15.81 15.67 15.11
OpenBenchmarking.org images/sec, More Is Better TensorFlow 2.12 Device: GPU - Batch Size: 32 - Model: ResNet-50 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 1.2533 2.5066 3.7599 5.0132 6.2665 SE +/- 0.01, N = 2 SE +/- 0.01, N = 2 SE +/- 0.02, N = 3 SE +/- 0.01, N = 3 SE +/- 0.00, N = 3 5.51 5.55 5.50 5.57 5.35
OpenBenchmarking.org images/sec, More Is Better TensorFlow 2.12 Device: GPU - Batch Size: 64 - Model: GoogLeNet NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 4 8 12 16 20 SE +/- 0.07, N = 3 SE +/- 0.06, N = 2 SE +/- 0.08, N = 3 SE +/- 0.09, N = 3 15.52 15.54 15.50 15.63 15.00
OpenBenchmarking.org images/sec, More Is Better TensorFlow 2.12 Device: GPU - Batch Size: 64 - Model: ResNet-50 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 1.2533 2.5066 3.7599 5.0132 6.2665 SE +/- 0.01, N = 2 SE +/- 0.00, N = 3 SE +/- 0.01, N = 2 SE +/- 0.01, N = 3 SE +/- 0.02, N = 3 5.55 5.55 5.53 5.57 5.33
PlaidML This test profile uses PlaidML deep learning framework developed by Intel for offering up various benchmarks. Learn more via the OpenBenchmarking.org test page.
FP16: No - Mode: Training - Network: Mobilenet - Device: OpenCL
NVIDIA RTX 4070 SUPER: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 4070: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test run did not produce a result.
NVIDIA RTX 4070 TI: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 3090: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 4070 TI SUPER: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
FP16: No - Mode: Inference - Network: IMDB LSTM - Device: OpenCL
NVIDIA RTX 4070 SUPER: The test run did not produce a result. The test run did not produce a result. The test quit with a non-zero exit status.
NVIDIA RTX 4070: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 4070 TI: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 3090: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 4070 TI SUPER: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
FP16: No - Mode: Inference - Network: Mobilenet - Device: OpenCL
NVIDIA RTX 4070 SUPER: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 4070: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 4070 TI: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 3090: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 4070 TI SUPER: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
FP16: Yes - Mode: Inference - Network: Mobilenet - Device: OpenCL
NVIDIA RTX 4070 SUPER: The test run did not produce a result. The test run did not produce a result. The test quit with a non-zero exit status. E: AttributeError: 'method_descriptor' object has no attribute 'default'
NVIDIA RTX 4070: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 4070 TI: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 3090: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 4070 TI SUPER: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
FP16: No - Mode: Inference - Network: DenseNet 201 - Device: OpenCL
NVIDIA RTX 4070 SUPER: The test run did not produce a result. The test quit with a non-zero exit status. The test run did not produce a result.
NVIDIA RTX 4070: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 4070 TI: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 3090: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 4070 TI SUPER: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
LeelaChessZero LeelaChessZero (lc0 / lczero) is a chess engine automated vian neural networks. This test profile can be used for OpenCL, CUDA + cuDNN, and BLAS (CPU-based) benchmarking. Learn more via the OpenBenchmarking.org test page.
Backend: OpenCL
NVIDIA RTX 4070 SUPER: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 4070: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 4070 TI: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 3090: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 4070 TI SUPER: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
PyTorch This is a benchmark of PyTorch making use of pytorch-benchmark [https://github.com/LukasHedegaard/pytorch-benchmark]. Currently this test profile is catered to CPU-based testing. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org batches/sec, More Is Better PyTorch 2.1 Device: NVIDIA CUDA GPU - Batch Size: 1 - Model: ResNet-50 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 120 240 360 480 600 SE +/- 3.09, N = 3 SE +/- 11.16, N = 12 SE +/- 3.07, N = 3 557.73 546.76 535.39 525.12 558.82 MIN: 513.63 / MAX: 563.37 MIN: 195.25 / MAX: 556.94 MIN: 428.43 / MAX: 572.99 MIN: 458.54 / MAX: 542.46 MIN: 473.77 / MAX: 573.46
OpenBenchmarking.org batches/sec, More Is Better PyTorch 2.1 Device: NVIDIA CUDA GPU - Batch Size: 1 - Model: ResNet-152 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 40 80 120 160 200 SE +/- 0.36, N = 3 SE +/- 0.73, N = 3 SE +/- 0.09, N = 2 SE +/- 0.38, N = 3 201.94 198.18 201.19 197.12 200.46 MIN: 183.53 / MAX: 206.5 MIN: 181.27 / MAX: 200.06 MIN: 180.79 / MAX: 203.92 MIN: 137.37 / MAX: 198.9 MIN: 177.25 / MAX: 203.31
OpenBenchmarking.org batches/sec, More Is Better PyTorch 2.1 Device: NVIDIA CUDA GPU - Batch Size: 16 - Model: ResNet-50 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 120 240 360 480 600 SE +/- 0.26, N = 3 SE +/- 2.23, N = 3 SE +/- 0.89, N = 2 SE +/- 1.33, N = 3 509.45 458.39 502.92 419.76 531.96 MIN: 430.1 / MAX: 516.48 MIN: 404.5 / MAX: 461.01 MIN: 415.65 / MAX: 520.39 MIN: 376.2 / MAX: 422.17 MIN: 422.98 / MAX: 539.81
OpenBenchmarking.org batches/sec, More Is Better PyTorch 2.1 Device: NVIDIA CUDA GPU - Batch Size: 32 - Model: ResNet-50 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 120 240 360 480 600 SE +/- 2.17, N = 2 SE +/- 0.13, N = 2 SE +/- 1.69, N = 3 SE +/- 0.70, N = 3 501.50 459.94 505.55 420.29 532.77 MIN: 415.94 / MAX: 510.69 MIN: 403.65 / MAX: 462.59 MIN: 419.93 / MAX: 512.69 MIN: 376.81 / MAX: 421.58 MIN: 420.31 / MAX: 538.98
OpenBenchmarking.org batches/sec, More Is Better PyTorch 2.1 Device: NVIDIA CUDA GPU - Batch Size: 64 - Model: ResNet-50 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 110 220 330 440 550 SE +/- 0.92, N = 3 SE +/- 0.27, N = 3 SE +/- 1.92, N = 3 SE +/- 0.24, N = 3 SE +/- 1.58, N = 3 507.45 458.36 505.62 419.03 527.82 MIN: 423.41 / MAX: 512.88 MIN: 404.89 / MAX: 461.01 MIN: 426.6 / MAX: 513.25 MIN: 376 / MAX: 422 MIN: 419.39 / MAX: 534.44
OpenBenchmarking.org batches/sec, More Is Better PyTorch 2.1 Device: NVIDIA CUDA GPU - Batch Size: 16 - Model: ResNet-152 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 40 80 120 160 200 SE +/- 0.29, N = 3 SE +/- 0.33, N = 3 195.40 187.26 194.29 164.14 198.58 MIN: 186.09 / MAX: 197.7 MIN: 179.81 / MAX: 188.21 MIN: 182.25 / MAX: 197.39 MIN: 145.67 / MAX: 165.38 MIN: 183.91 / MAX: 201.98
OpenBenchmarking.org batches/sec, More Is Better PyTorch 2.1 Device: NVIDIA CUDA GPU - Batch Size: 256 - Model: ResNet-50 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 110 220 330 440 550 SE +/- 1.39, N = 3 SE +/- 0.34, N = 3 SE +/- 0.14, N = 3 SE +/- 0.54, N = 3 504.67 459.93 416.89 529.14 MIN: 412.34 / MAX: 514.07 MIN: 403.65 / MAX: 462.74 MIN: 329.77 / MAX: 420.82 MIN: 414.54 / MAX: 534.65
Device: NVIDIA CUDA GPU - Batch Size: 256 - Model: ResNet-50
NVIDIA RTX 4070 TI: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: TypeError: 'NoneType' object is not callable
OpenBenchmarking.org batches/sec, More Is Better PyTorch 2.1 Device: NVIDIA CUDA GPU - Batch Size: 32 - Model: ResNet-152 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 40 80 120 160 200 SE +/- 0.29, N = 3 SE +/- 0.28, N = 3 195.39 187.69 198.82 163.74 197.82 MIN: 183.94 / MAX: 198.7 MIN: 182.03 / MAX: 188.31 MIN: 188.33 / MAX: 201.47 MIN: 144.93 / MAX: 165.03 MIN: 176.19 / MAX: 201.63
OpenBenchmarking.org batches/sec, More Is Better PyTorch 2.1 Device: NVIDIA CUDA GPU - Batch Size: 512 - Model: ResNet-50 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 110 220 330 440 550 SE +/- 4.43, N = 2 SE +/- 0.43, N = 2 SE +/- 0.83, N = 2 SE +/- 0.40, N = 3 SE +/- 1.16, N = 3 504.27 459.27 504.66 416.20 529.49 MIN: 418.22 / MAX: 512.44 MIN: 405.48 / MAX: 461.88 MIN: 424.27 / MAX: 509.08 MIN: 355.45 / MAX: 419.05 MIN: 410.12 / MAX: 537.25
OpenBenchmarking.org batches/sec, More Is Better PyTorch 2.1 Device: NVIDIA CUDA GPU - Batch Size: 64 - Model: ResNet-152 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 40 80 120 160 200 SE +/- 0.51, N = 3 SE +/- 0.34, N = 3 SE +/- 0.78, N = 2 SE +/- 0.20, N = 3 196.07 186.63 197.02 164.14 196.50 MIN: 171.95 / MAX: 199.96 MIN: 180.51 / MAX: 187.79 MIN: 183.92 / MAX: 200.54 MIN: 149 / MAX: 165 MIN: 179.34 / MAX: 200
OpenBenchmarking.org batches/sec, More Is Better PyTorch 2.1 Device: NVIDIA CUDA GPU - Batch Size: 256 - Model: ResNet-152 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 40 80 120 160 200 SE +/- 1.14, N = 2 SE +/- 0.17, N = 3 SE +/- 0.19, N = 2 SE +/- 0.95, N = 3 194.58 187.27 195.86 161.01 198.70 MIN: 183.74 / MAX: 198.52 MIN: 179.9 / MAX: 188.08 MIN: 181.64 / MAX: 199.2 MIN: 138.12 / MAX: 165.16 MIN: 185.21 / MAX: 203.36
OpenBenchmarking.org batches/sec, More Is Better PyTorch 2.1 Device: NVIDIA CUDA GPU - Batch Size: 512 - Model: ResNet-152 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 40 80 120 160 200 SE +/- 1.38, N = 2 SE +/- 0.05, N = 3 SE +/- 0.33, N = 2 SE +/- 0.81, N = 3 195.30 187.51 194.87 164.35 198.01 MIN: 182 / MAX: 199.43 MIN: 181.57 / MAX: 188.05 MIN: 180.8 / MAX: 198 MIN: 149.91 / MAX: 166.09 MIN: 185.3 / MAX: 202.59
OpenBenchmarking.org batches/sec, More Is Better PyTorch 2.1 Device: NVIDIA CUDA GPU - Batch Size: 1 - Model: Efficientnet_v2_l NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 20 40 60 80 100 SE +/- 0.55, N = 3 SE +/- 0.33, N = 3 SE +/- 0.24, N = 2 106.37 107.59 108.59 105.55 105.86 MIN: 97.91 / MAX: 108.16 MIN: 98.77 / MAX: 109.43 MIN: 99.04 / MAX: 110.68 MIN: 91.76 / MAX: 107.42 MIN: 95.05 / MAX: 107.6
OpenBenchmarking.org batches/sec, More Is Better PyTorch 2.1 Device: NVIDIA CUDA GPU - Batch Size: 16 - Model: Efficientnet_v2_l NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 20 40 60 80 100 SE +/- 0.52, N = 2 SE +/- 0.53, N = 2 SE +/- 0.33, N = 3 103.68 103.45 98.11 103.66 MIN: 96.86 / MAX: 105.56 MIN: 95.22 / MAX: 105.88 MIN: 89.88 / MAX: 100.25 MIN: 93.46 / MAX: 105.95
Device: NVIDIA CUDA GPU - Batch Size: 16 - Model: Efficientnet_v2_l
NVIDIA RTX 4070 SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: AttributeError: 'tuple' object has no attribute '_compiled_call_impl'
OpenBenchmarking.org batches/sec, More Is Better PyTorch 2.1 Device: NVIDIA CUDA GPU - Batch Size: 32 - Model: Efficientnet_v2_l NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 20 40 60 80 100 SE +/- 6.65, N = 5 SE +/- 0.13, N = 3 SE +/- 0.62, N = 3 102.60 102.90 96.50 99.05 102.83 MIN: 94.84 / MAX: 104.25 MIN: 95.98 / MAX: 104.54 MIN: 64.35 / MAX: 104.79 MIN: 91.8 / MAX: 100.69 MIN: 92.44 / MAX: 105.47
OpenBenchmarking.org batches/sec, More Is Better PyTorch 2.1 Device: NVIDIA CUDA GPU - Batch Size: 64 - Model: Efficientnet_v2_l NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 20 40 60 80 100 SE +/- 1.49, N = 2 SE +/- 0.45, N = 3 SE +/- 0.39, N = 2 SE +/- 0.14, N = 3 SE +/- 0.13, N = 3 102.60 101.55 103.20 99.84 103.49 MIN: 79.69 / MAX: 105.28 MIN: 93.44 / MAX: 103.08 MIN: 95.31 / MAX: 105.27 MIN: 92.73 / MAX: 101.46 MIN: 93.23 / MAX: 105.43
OpenBenchmarking.org batches/sec, More Is Better PyTorch 2.1 Device: NVIDIA CUDA GPU - Batch Size: 256 - Model: Efficientnet_v2_l NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 20 40 60 80 100 SE +/- 0.05, N = 2 SE +/- 0.57, N = 3 SE +/- 0.18, N = 3 103.17 101.24 103.24 99.43 102.83 MIN: 95.79 / MAX: 105.15 MIN: 93.33 / MAX: 102.92 MIN: 95.41 / MAX: 104.9 MIN: 90.49 / MAX: 101.97 MIN: 93.16 / MAX: 105.07
OpenBenchmarking.org batches/sec, More Is Better PyTorch 2.1 Device: NVIDIA CUDA GPU - Batch Size: 512 - Model: Efficientnet_v2_l NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 20 40 60 80 100 SE +/- 0.39, N = 3 SE +/- 0.36, N = 2 SE +/- 0.19, N = 3 103.57 101.43 103.50 99.25 103.53 MIN: 95.95 / MAX: 105.54 MIN: 93.27 / MAX: 103.58 MIN: 94.95 / MAX: 105.61 MIN: 91.16 / MAX: 101.18 MIN: 88.81 / MAX: 104.8
Caffe This is a benchmark of the Caffe deep learning framework and currently supports the AlexNet and Googlenet model and execution on both CPUs and NVIDIA GPUs. Learn more via the OpenBenchmarking.org test page.
Model: AlexNet - Acceleration: NVIDIA CUDA - Iterations: 100
NVIDIA RTX 4070 SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7dd7c6de3450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 4070: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7b80311e3450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 4070 TI: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x736df4b59450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 3090: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x77ed97de3450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 4070 TI SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x78750ffea450 google::LogMessageFatal::~LogMessageFatal()
Model: AlexNet - Acceleration: NVIDIA CUDA - Iterations: 200
NVIDIA RTX 4070 SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7b5ea59be450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 4070: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7c31ed79d450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 4070 TI: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7ba579075450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 3090: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7ace5f7b4450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 4070 TI SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x70294d9e3450 google::LogMessageFatal::~LogMessageFatal()
Model: AlexNet - Acceleration: NVIDIA CUDA - Iterations: 1000
NVIDIA RTX 4070 SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7670bcda4450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 4070: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7bb89c5be450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 4070 TI: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x72248ee5c450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 3090: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7d66735f5450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 4070 TI SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7ce671f7d450 google::LogMessageFatal::~LogMessageFatal()
Model: GoogleNet - Acceleration: NVIDIA CUDA - Iterations: 100
NVIDIA RTX 4070 SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x73552c3e3450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 4070: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x71f0ea05a450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 4070 TI: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7898abd73450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 3090: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7522f0d76450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 4070 TI SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x703a2cf99450 google::LogMessageFatal::~LogMessageFatal()
Model: GoogleNet - Acceleration: NVIDIA CUDA - Iterations: 200
NVIDIA RTX 4070 SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7d7151816450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 4070: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7e64df79d450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 4070 TI: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x761e63d48450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 3090: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7bfcc77e3450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 4070 TI SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7bc837b4a450 google::LogMessageFatal::~LogMessageFatal()
Model: GoogleNet - Acceleration: NVIDIA CUDA - Iterations: 1000
NVIDIA RTX 4070 SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x74746a490450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 4070: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7493bdbbc450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 4070 TI: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7338f7773450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 3090: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7792f141e450 google::LogMessageFatal::~LogMessageFatal()
NVIDIA RTX 4070 TI SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: @ 0x7911d0e44450 google::LogMessageFatal::~LogMessageFatal()
NCNN NCNN is a high performance neural network inference framework optimized for mobile and other platforms developed by Tencent. Learn more via the OpenBenchmarking.org test page.
Target: Vulkan GPU
NVIDIA RTX 4070 SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: ncnn: line 3: ./benchncnn: No such file or directory
OpenBenchmarking.org ms, Fewer Is Better NCNN 20230517 Target: Vulkan GPU - Model: mobilenet NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER 3 6 9 12 15 SE +/- 0.21, N = 12 SE +/- 0.25, N = 12 SE +/- 0.22, N = 9 SE +/- 0.22, N = 12 SE +/- 0.47, N = 9 7.20 7.45 6.92 7.48 8.62 MIN: 6.2 / MAX: 11.13 MIN: 6.87 / MAX: 734.65 MIN: 6.06 / MAX: 8.65 MIN: 6.44 / MAX: 11.44 MIN: 6.42 / MAX: 1101.3 1. (CXX) g++ options: -O3 -rdynamic -lgomp -lpthread
OpenBenchmarking.org ms, Fewer Is Better NCNN 20230517 Target: Vulkan GPU-v2-v2 - Model: mobilenet-v2 NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER 1.0553 2.1106 3.1659 4.2212 5.2765 SE +/- 0.07, N = 12 SE +/- 0.06, N = 12 SE +/- 0.05, N = 9 SE +/- 0.03, N = 12 SE +/- 0.44, N = 9 2.48 2.54 2.67 2.70 3.03 MIN: 2.02 / MAX: 5.82 MIN: 2.07 / MAX: 15.69 MIN: 2.44 / MAX: 22.79 MIN: 2.42 / MAX: 5.99 MIN: 2.38 / MAX: 970.87 1. (CXX) g++ options: -O3 -rdynamic -lgomp -lpthread
OpenBenchmarking.org ms, Fewer Is Better NCNN 20230517 Target: Vulkan GPU-v3-v3 - Model: mobilenet-v3 NVIDIA RTX 4070 NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 TI 2 4 6 8 10 SE +/- 0.08, N = 12 SE +/- 0.09, N = 9 SE +/- 0.08, N = 12 SE +/- 0.16, N = 9 SE +/- 0.09, N = 9 2.15 2.20 2.16 2.25 2.09 MIN: 1.81 / MAX: 2.58 MIN: 1.91 / MAX: 2.71 MIN: 1.78 / MAX: 18.62 MIN: 1.75 / MAX: 343.7 MIN: 1.78 / MAX: 2.85 1. (CXX) g++ options: -O3 -rdynamic -lgomp -lpthread
OpenBenchmarking.org ms, Fewer Is Better NCNN 20230517 Target: Vulkan GPU - Model: shufflenet-v2 NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER 0.9383 1.8766 2.8149 3.7532 4.6915 SE +/- 0.09, N = 11 SE +/- 0.08, N = 12 SE +/- 0.10, N = 9 SE +/- 0.09, N = 11 SE +/- 0.34, N = 8 2.08 2.01 2.09 2.13 2.31 MIN: 1.82 / MAX: 2.59 MIN: 1.73 / MAX: 3.86 MIN: 1.83 / MAX: 2.96 MIN: 1.82 / MAX: 3.66 MIN: 1.76 / MAX: 421.42 1. (CXX) g++ options: -O3 -rdynamic -lgomp -lpthread
OpenBenchmarking.org ms, Fewer Is Better NCNN 20230517 Target: Vulkan GPU - Model: mnasnet NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER 0.9315 1.863 2.7945 3.726 4.6575 SE +/- 0.05, N = 10 SE +/- 1.94, N = 12 SE +/- 0.06, N = 9 SE +/- 0.06, N = 12 SE +/- 1.31, N = 9 2.24 4.14 2.30 2.26 3.85 MIN: 2 / MAX: 2.6 MIN: 1.95 / MAX: 1020 MIN: 2.01 / MAX: 2.64 MIN: 1.99 / MAX: 4.41 MIN: 1.89 / MAX: 1093.29 1. (CXX) g++ options: -O3 -rdynamic -lgomp -lpthread
OpenBenchmarking.org ms, Fewer Is Better NCNN 20230517 Target: Vulkan GPU - Model: efficientnet-b0 NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER 4 8 12 16 20 SE +/- 0.20, N = 12 SE +/- 0.07, N = 12 SE +/- 0.07, N = 9 SE +/- 0.07, N = 12 SE +/- 0.97, N = 9 3.59 3.46 3.54 3.48 5.07 MIN: 3.02 / MAX: 509.97 MIN: 3.13 / MAX: 7.03 MIN: 3.19 / MAX: 3.95 MIN: 3.18 / MAX: 5.61 MIN: 3.22 / MAX: 1124.2 1. (CXX) g++ options: -O3 -rdynamic -lgomp -lpthread
OpenBenchmarking.org ms, Fewer Is Better NCNN 20230517 Target: Vulkan GPU - Model: blazeface NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 0.198 0.396 0.594 0.792 0.99 SE +/- 0.03, N = 12 SE +/- 0.03, N = 9 SE +/- 0.02, N = 12 SE +/- 0.04, N = 9 SE +/- 0.03, N = 9 0.82 0.84 0.88 0.84 0.84 MIN: 0.63 / MAX: 2.58 MIN: 0.63 / MAX: 1.13 MIN: 0.74 / MAX: 2.72 MIN: 0.65 / MAX: 4.63 MIN: 0.64 / MAX: 0.96 1. (CXX) g++ options: -O3 -rdynamic -lgomp -lpthread
OpenBenchmarking.org ms, Fewer Is Better NCNN 20230517 Target: Vulkan GPU - Model: googlenet NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 3 6 9 12 15 SE +/- 1.04, N = 12 SE +/- 0.18, N = 9 SE +/- 0.12, N = 12 SE +/- 1.21, N = 9 SE +/- 0.14, N = 9 7.37 6.11 6.46 11.04 6.06 MIN: 5.33 / MAX: 1653.33 MIN: 5.25 / MAX: 9.16 MIN: 5.71 / MAX: 8.81 MIN: 5.28 / MAX: 1769.19 MIN: 5.33 / MAX: 8.36 1. (CXX) g++ options: -O3 -rdynamic -lgomp -lpthread
OpenBenchmarking.org ms, Fewer Is Better NCNN 20230517 Target: Vulkan GPU - Model: vgg16 NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER 30 60 90 120 150 SE +/- 13.24, N = 12 SE +/- 8.53, N = 12 SE +/- 5.69, N = 9 SE +/- 2.41, N = 12 SE +/- 29.60, N = 9 45.52 34.49 24.45 24.85 117.81 MIN: 17.49 / MAX: 643.35 MIN: 17.35 / MAX: 644.35 MIN: 18.11 / MAX: 642.2 MIN: 21.65 / MAX: 648.84 MIN: 17.16 / MAX: 647.67 1. (CXX) g++ options: -O3 -rdynamic -lgomp -lpthread
OpenBenchmarking.org ms, Fewer Is Better NCNN 20230517 Target: Vulkan GPU - Model: resnet18 NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER 4 8 12 16 20 SE +/- 0.73, N = 12 SE +/- 1.74, N = 12 SE +/- 3.28, N = 9 SE +/- 1.66, N = 12 SE +/- 3.49, N = 9 5.11 7.74 8.94 7.58 8.97 MIN: 3.99 / MAX: 916.69 MIN: 3.96 / MAX: 898.97 MIN: 4.04 / MAX: 900 MIN: 4.36 / MAX: 907.95 MIN: 3.94 / MAX: 922.04 1. (CXX) g++ options: -O3 -rdynamic -lgomp -lpthread
OpenBenchmarking.org ms, Fewer Is Better NCNN 20230517 Target: Vulkan GPU - Model: alexnet NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER 4 8 12 16 20 SE +/- 1.70, N = 12 SE +/- 2.33, N = 12 SE +/- 1.74, N = 9 SE +/- 0.02, N = 12 SE +/- 5.86, N = 9 5.78 6.07 6.20 4.41 16.17 MIN: 3.6 / MAX: 397.75 MIN: 3.63 / MAX: 432.24 MIN: 3.56 / MAX: 340 MIN: 4.29 / MAX: 4.98 MIN: 3.52 / MAX: 436.52 1. (CXX) g++ options: -O3 -rdynamic -lgomp -lpthread
OpenBenchmarking.org ms, Fewer Is Better NCNN 20230517 Target: Vulkan GPU - Model: resnet50 NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER 10 20 30 40 50 SE +/- 0.49, N = 12 SE +/- 4.00, N = 12 SE +/- 0.12, N = 9 SE +/- 0.06, N = 12 SE +/- 14.70, N = 9 8.72 12.25 8.20 8.79 46.26 MIN: 7.94 / MAX: 935.44 MIN: 8 / MAX: 1777.17 MIN: 7.69 / MAX: 11.69 MIN: 8.45 / MAX: 10.96 MIN: 7.71 / MAX: 1829.99 1. (CXX) g++ options: -O3 -rdynamic -lgomp -lpthread
OpenBenchmarking.org ms, Fewer Is Better NCNN 20230517 Target: Vulkan GPU - Model: yolov4-tiny NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER 14 28 42 56 70 SE +/- 5.37, N = 12 SE +/- 3.10, N = 12 SE +/- 2.52, N = 9 SE +/- 4.57, N = 12 SE +/- 10.56, N = 9 20.74 16.37 13.31 17.20 63.82 MIN: 10.3 / MAX: 854.36 MIN: 10.57 / MAX: 855.36 MIN: 10.13 / MAX: 840.81 MIN: 11.23 / MAX: 855.47 MIN: 10.28 / MAX: 858.44 1. (CXX) g++ options: -O3 -rdynamic -lgomp -lpthread
OpenBenchmarking.org ms, Fewer Is Better NCNN 20230517 Target: Vulkan GPU - Model: squeezenet_ssd NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER 2 4 6 8 10 SE +/- 0.12, N = 12 SE +/- 0.93, N = 12 SE +/- 0.20, N = 9 SE +/- 0.13, N = 12 SE +/- 1.76, N = 9 5.18 6.13 5.20 5.19 6.86 MIN: 4.67 / MAX: 6.88 MIN: 4.75 / MAX: 1633.33 MIN: 4.36 / MAX: 6.17 MIN: 4.65 / MAX: 14.18 MIN: 4.34 / MAX: 1630.01 1. (CXX) g++ options: -O3 -rdynamic -lgomp -lpthread
OpenBenchmarking.org ms, Fewer Is Better NCNN 20230517 Target: Vulkan GPU - Model: regnety_400m NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER 3 6 9 12 15 SE +/- 0.24, N = 12 SE +/- 0.18, N = 12 SE +/- 0.32, N = 8 SE +/- 0.26, N = 12 SE +/- 3.28, N = 9 6.21 5.89 6.47 6.59 11.11 MIN: 5.53 / MAX: 8.99 MIN: 5.42 / MAX: 7.57 MIN: 5.44 / MAX: 9.3 MIN: 5.45 / MAX: 9.09 MIN: 5.49 / MAX: 4942.19 1. (CXX) g++ options: -O3 -rdynamic -lgomp -lpthread
OpenBenchmarking.org ms, Fewer Is Better NCNN 20230517 Target: Vulkan GPU - Model: vision_transformer NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER 200 400 600 800 1000 SE +/- 58.37, N = 12 SE +/- 58.07, N = 12 SE +/- 52.80, N = 9 SE +/- 57.46, N = 12 SE +/- 87.53, N = 9 382.82 497.66 327.82 312.10 844.61 MIN: 46.51 / MAX: 1840 MIN: 46.41 / MAX: 1886.67 MIN: 46.48 / MAX: 1816.93 MIN: 47.85 / MAX: 1850.09 MIN: 46.34 / MAX: 1866.93 1. (CXX) g++ options: -O3 -rdynamic -lgomp -lpthread
OpenBenchmarking.org ms, Fewer Is Better NCNN 20230517 Target: Vulkan GPU - Model: FastestDet NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER 2 4 6 8 10 SE +/- 0.13, N = 12 SE +/- 0.32, N = 10 SE +/- 0.08, N = 8 SE +/- 0.11, N = 12 SE +/- 0.29, N = 9 2.67 3.04 2.50 2.55 2.86 MIN: 2.04 / MAX: 3.47 MIN: 2.14 / MAX: 847.58 MIN: 2.1 / MAX: 32.36 MIN: 2.08 / MAX: 5.83 MIN: 2.17 / MAX: 577.17 1. (CXX) g++ options: -O3 -rdynamic -lgomp -lpthread
Rodinia Rodinia is a suite focused upon accelerating compute-intensive applications with accelerators. CUDA, OpenMP, and OpenCL parallel models are supported by the included applications. This profile utilizes select OpenCL, NVIDIA CUDA and OpenMP test binaries at the moment. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org Seconds, Fewer Is Better Rodinia 3.1 Test: OpenCL Particle Filter NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 0.9221 1.8442 2.7663 3.6884 4.6105 SE +/- 0.039, N = 4 SE +/- 0.008, N = 3 SE +/- 0.002, N = 3 SE +/- 0.030, N = 15 SE +/- 0.004, N = 3 3.480 4.098 3.291 3.844 2.973 1. (CXX) g++ options: -O2 -lOpenCL
ArrayFire ArrayFire is an GPU and CPU numeric processing library, this test uses the built-in CPU and OpenCL ArrayFire benchmarks. Learn more via the OpenBenchmarking.org test page.
Test: Conjugate Gradient OpenCL
NVIDIA RTX 4070 SUPER: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result. E: arrayfire: line 3: ./cg_opencl: No such file or directory
NVIDIA RTX 4070: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result. E: arrayfire: line 3: ./cg_opencl: No such file or directory
NVIDIA RTX 4070 TI: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result. E: arrayfire: line 3: ./cg_opencl: No such file or directory
NVIDIA RTX 3090: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result. E: arrayfire: line 3: ./cg_opencl: No such file or directory
NVIDIA RTX 4070 TI SUPER: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result. E: arrayfire: line 3: ./cg_opencl: No such file or directory
ProjectPhysX OpenCL-Benchmark ProjectPhysX OpenCL-Benchmark provides various OpenCL compute and memory bandwidth micro-benchmarks Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org TFLOPs/s, More Is Better ProjectPhysX OpenCL-Benchmark 1.2 Operation: FP32 Compute NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 10 20 30 40 50 SE +/- 0.03, N = 3 SE +/- 0.03, N = 3 SE +/- 0.00, N = 3 SE +/- 0.10, N = 3 SE +/- 0.01, N = 3 38.59 31.77 40.91 39.40 45.95 1. (CXX) g++ options: -std=c++17 -pthread -lOpenCL
Blender Blender is an open-source 3D creation and modeling software project. This test is of Blender's Cycles performance with various sample files. GPU computing via NVIDIA OptiX and NVIDIA CUDA is currently supported as well as HIP for AMD Radeon GPUs and Intel oneAPI for Intel Graphics. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org Seconds, Fewer Is Better Blender 4.0 Blend File: BMW27 - Compute: NVIDIA OptiX NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 2 4 6 8 10 SE +/- 0.06, N = 13 SE +/- 0.01, N = 3 SE +/- 0.02, N = 3 SE +/- 0.06, N = 14 SE +/- 0.06, N = 14 5.57 6.21 5.43 6.31 5.04
OpenBenchmarking.org Seconds, Fewer Is Better Blender 4.0 Blend File: Classroom - Compute: NVIDIA OptiX NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 4 8 12 16 20 SE +/- 0.00, N = 3 SE +/- 0.03, N = 3 SE +/- 0.03, N = 3 SE +/- 0.02, N = 3 SE +/- 0.04, N = 3 12.60 14.86 12.30 15.26 11.20
OpenBenchmarking.org Seconds, Fewer Is Better Blender 4.0 Blend File: Fishy Cat - Compute: NVIDIA OptiX NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 3 6 9 12 15 SE +/- 0.06, N = 13 SE +/- 0.03, N = 3 SE +/- 0.01, N = 3 SE +/- 0.08, N = 9 SE +/- 0.06, N = 13 9.45 11.03 9.02 10.64 8.32
OpenBenchmarking.org Seconds, Fewer Is Better Blender 4.0 Blend File: Barbershop - Compute: NVIDIA OptiX NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 13 26 39 52 65 SE +/- 0.10, N = 3 SE +/- 0.04, N = 3 SE +/- 0.05, N = 3 SE +/- 0.02, N = 2 SE +/- 0.08, N = 3 51.30 58.44 50.73 54.30 44.49
OpenBenchmarking.org Seconds, Fewer Is Better Blender 4.0 Blend File: Pabellon Barcelona - Compute: NVIDIA OptiX NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 4 8 12 16 20 SE +/- 0.03, N = 3 SE +/- 0.01, N = 3 SE +/- 0.01, N = 3 SE +/- 0.02, N = 3 SE +/- 0.02, N = 3 14.29 16.55 13.97 17.30 12.56
NeatBench NeatBench is a benchmark of the cross-platform Neat Video software on the CPU and optional GPU (OpenCL / CUDA) support. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org FPS, More Is Better NeatBench 5 Acceleration: GPU NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 900 1800 2700 3600 4500 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 512.75, N = 16 4070.0 4070.0 4070.0 3090.0 2084.1
IndigoBench This is a test of Indigo Renderer's IndigoBench benchmark. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org M samples/s, More Is Better IndigoBench 4.4 Acceleration: OpenCL GPU - Scene: Bedroom NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 6 12 18 24 30 SE +/- 0.01, N = 3 SE +/- 0.01, N = 3 SE +/- 0.03, N = 3 SE +/- 0.01, N = 3 SE +/- 0.01, N = 3 19.80 18.20 20.26 20.96 24.57
OpenBenchmarking.org M samples/s, More Is Better IndigoBench 4.4 Acceleration: OpenCL GPU - Scene: Supercar NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 14 28 42 56 70 SE +/- 0.03, N = 3 SE +/- 0.03, N = 3 SE +/- 0.02, N = 3 SE +/- 0.03, N = 3 SE +/- 0.06, N = 3 52.81 48.52 53.59 52.01 61.34
LuxCoreRender LuxCoreRender is an open-source 3D physically based renderer formerly known as LuxRender. LuxCoreRender supports CPU-based rendering as well as GPU acceleration via OpenCL, NVIDIA CUDA, and NVIDIA OptiX interfaces. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org M samples/sec, More Is Better LuxCoreRender 2.6 Scene: DLSC - Acceleration: GPU NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 4 8 12 16 20 SE +/- 0.01, N = 3 SE +/- 0.01, N = 3 SE +/- 0.00, N = 3 SE +/- 1.13, N = 12 SE +/- 0.01, N = 3 13.59 11.74 13.95 12.99 16.23 MIN: 12.52 / MAX: 13.84 MIN: 11.35 / MAX: 11.83 MIN: 13.67 / MAX: 14.14 MIN: 0.52 / MAX: 14.69 MIN: 15.91 / MAX: 16.36
OpenBenchmarking.org M samples/sec, More Is Better LuxCoreRender 2.6 Scene: Danish Mood - Acceleration: GPU NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 3 6 9 12 15 SE +/- 0.08, N = 3 SE +/- 0.06, N = 3 SE +/- 0.11, N = 3 SE +/- 0.04, N = 3 SE +/- 0.03, N = 3 10.56 8.89 10.99 10.20 12.42 MIN: 3.7 / MAX: 12.17 MIN: 3.32 / MAX: 10.26 MIN: 4.17 / MAX: 12.71 MIN: 4.07 / MAX: 11.93 MIN: 4.35 / MAX: 14.32
OpenBenchmarking.org M samples/sec, More Is Better LuxCoreRender 2.6 Scene: Orange Juice - Acceleration: GPU NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 4 8 12 16 20 SE +/- 0.07, N = 3 SE +/- 0.03, N = 3 SE +/- 0.00, N = 3 SE +/- 0.01, N = 3 SE +/- 0.15, N = 4 11.72 10.40 11.89 12.14 13.64 MIN: 9.6 / MAX: 15.44 MIN: 8.31 / MAX: 13.9 MIN: 9.85 / MAX: 15.88 MIN: 10.24 / MAX: 16.71 MIN: 11.16 / MAX: 18.46
OpenBenchmarking.org M samples/sec, More Is Better LuxCoreRender 2.6 Scene: LuxCore Benchmark - Acceleration: GPU NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 4 8 12 16 20 SE +/- 0.02, N = 3 SE +/- 0.01, N = 3 SE +/- 0.01, N = 3 SE +/- 0.03, N = 2 SE +/- 0.00, N = 3 12.82 10.92 13.23 13.12 14.61 MIN: 4.84 / MAX: 14.62 MIN: 4.45 / MAX: 12.42 MIN: 5.41 / MAX: 15.13 MIN: 4.85 / MAX: 15.21 MIN: 5.91 / MAX: 16.88
OpenBenchmarking.org M samples/sec, More Is Better LuxCoreRender 2.6 Scene: Rainbow Colors and Prism - Acceleration: GPU NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 8 16 24 32 40 SE +/- 0.03, N = 3 SE +/- 0.01, N = 3 SE +/- 0.07, N = 3 SE +/- 0.36, N = 5 SE +/- 0.09, N = 3 27.67 23.26 27.71 33.29 31.86 MIN: 24.87 / MAX: 29.03 MIN: 20.92 / MAX: 24.3 MIN: 25.01 / MAX: 29.15 MIN: 30.4 / MAX: 36.21 MIN: 28.57 / MAX: 33.29
Hashcat Hashcat is an open-source, advanced password recovery tool supporting GPU acceleration with OpenCL, NVIDIA CUDA, and Radeon ROCm. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org H/s, More Is Better Hashcat 6.2.4 Benchmark: MD5 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 20000M 40000M 60000M 80000M 100000M SE +/- 22430807.19, N = 3 SE +/- 33772046.30, N = 3 SE +/- 11283665.68, N = 3 SE +/- 53667246.37, N = 3 SE +/- 97655010.68, N = 3 67583033333 56147866667 73312233333 67177300000 82004966667
OpenBenchmarking.org H/s, More Is Better Hashcat 6.2.4 Benchmark: SHA1 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 6000M 12000M 18000M 24000M 30000M SE +/- 5140363.15, N = 3 SE +/- 6318315.53, N = 3 SE +/- 15926811.78, N = 3 SE +/- 26244639.66, N = 3 SE +/- 29067564.97, N = 3 22132600000 18202466667 23532400000 21323733333 26388600000
OpenBenchmarking.org H/s, More Is Better Hashcat 6.2.4 Benchmark: 7-Zip NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 300K 600K 900K 1200K 1500K SE +/- 1991.93, N = 3 SE +/- 2062.63, N = 3 SE +/- 2339.04, N = 3 SE +/- 1587.45, N = 3 SE +/- 1628.91, N = 3 1176467 976967 1262633 1056000 1420700
OpenBenchmarking.org H/s, More Is Better Hashcat 6.2.4 Benchmark: SHA-512 NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 800M 1600M 2400M 3200M 4000M SE +/- 1530068.99, N = 3 SE +/- 1059874.21, N = 3 SE +/- 721110.26, N = 3 SE +/- 3288532.26, N = 3 SE +/- 1098989.43, N = 3 3232733333 2673300000 3462500000 3081866667 3887033333
OpenBenchmarking.org H/s, More Is Better Hashcat 6.2.4 Benchmark: TrueCrypt RIPEMD160 + XTS NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 200K 400K 600K 800K 1000K SE +/- 633.33, N = 3 SE +/- 176.38, N = 3 SE +/- 888.82, N = 3 SE +/- 1757.21, N = 3 SE +/- 392.99, N = 3 802967 660967 858600 797833 961733
NAMD CUDA NAMD is a parallel molecular dynamics code designed for high-performance simulation of large biomolecular systems. NAMD was developed by the Theoretical and Computational Biophysics Group in the Beckman Institute for Advanced Science and Technology at the University of Illinois at Urbana-Champaign. This version of the NAMD test profile uses CUDA GPU acceleration. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org days/ns, Fewer Is Better NAMD CUDA 2.14 ATPase Simulation - 327,506 Atoms NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 0.0243 0.0486 0.0729 0.0972 0.1215 SE +/- 0.00031, N = 3 SE +/- 0.00021, N = 3 SE +/- 0.00061, N = 3 SE +/- 0.00042, N = 3 SE +/- 0.00018, N = 3 0.06791 0.07498 0.06788 0.10822 0.07715
FinanceBench FinanceBench is a collection of financial program benchmarks with support for benchmarking on the GPU via OpenCL and CPU benchmarking with OpenMP. The FinanceBench test cases are focused on Black-Sholes-Merton Process with Analytic European Option engine, QMC (Sobol) Monte-Carlo method (Equity Option Example), Bonds Fixed-rate bond with flat forward curve, and Repo Securities repurchase agreement. FinanceBench was originally written by the Cavazos Lab at University of Delaware. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org ms, Fewer Is Better FinanceBench 2016-07-25 Benchmark: Black-Scholes OpenCL NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 2 4 6 8 10 SE +/- 0.114, N = 15 SE +/- 0.003, N = 3 SE +/- 0.003, N = 3 SE +/- 0.006, N = 3 SE +/- 0.000, N = 3 5.912 6.906 5.226 5.741 0.501 1. (CXX) g++ options: -O3 -march=native -fopenmp
OpenBenchmarking.org GB/s, More Is Better cl-mem 2017-01-13 Benchmark: Read NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 200 400 600 800 1000 SE +/- 0.12, N = 3 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.32, N = 3 SE +/- 0.00, N = 3 446.2 446.3 446.3 825.8 595.2 1. (CC) gcc options: -O2 -flto -lOpenCL
OpenBenchmarking.org GB/s, More Is Better cl-mem 2017-01-13 Benchmark: Write NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 160 320 480 640 800 SE +/- 1.11, N = 3 SE +/- 0.55, N = 3 SE +/- 0.12, N = 3 SE +/- 0.83, N = 3 SE +/- 0.25, N = 3 407.5 406.7 412.2 753.8 551.9 1. (CC) gcc options: -O2 -flto -lOpenCL
clpeak Clpeak is designed to test the peak capabilities of OpenCL devices. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org GIOPS, More Is Better clpeak 1.1.2 OpenCL Test: Integer Compute INT NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 5K 10K 15K 20K 25K SE +/- 3.14, N = 3 SE +/- 15.26, N = 3 SE +/- 2.50, N = 3 SE +/- 16.49, N = 3 SE +/- 28.14, N = 3 18170.54 14555.19 19821.10 17923.33 22171.25 1. (CXX) g++ options: -O3
OpenBenchmarking.org GFLOPS, More Is Better clpeak 1.1.2 OpenCL Test: Single-Precision Float NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 9K 18K 27K 36K 45K SE +/- 0.99, N = 3 SE +/- 5.46, N = 3 SE +/- 11.67, N = 3 SE +/- 113.39, N = 3 SE +/- 50.25, N = 3 35492.69 28479.39 38691.73 34906.79 43244.79 1. (CXX) g++ options: -O3
OpenBenchmarking.org GFLOPS, More Is Better clpeak 1.1.2 OpenCL Test: Double-Precision Double NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 160 320 480 640 800 SE +/- 0.98, N = 3 SE +/- 0.21, N = 3 SE +/- 1.33, N = 3 SE +/- 1.63, N = 3 SE +/- 1.26, N = 3 630.11 515.17 667.05 642.23 750.36 1. (CXX) g++ options: -O3
OpenBenchmarking.org GBPS, More Is Better clpeak 1.1.2 OpenCL Test: Global Memory Bandwidth NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 200 400 600 800 1000 SE +/- 0.02, N = 3 SE +/- 0.02, N = 3 SE +/- 0.02, N = 3 SE +/- 0.02, N = 3 SE +/- 0.01, N = 3 437.65 437.21 437.63 816.55 582.84 1. (CXX) g++ options: -O3
MandelGPU MandelGPU is an OpenCL benchmark and this test runs with the OpenCL rendering float4 kernel with a maximum of 4096 iterations. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org Samples/sec, More Is Better MandelGPU 1.3pts1 OpenCL Device: GPU NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 140M 280M 420M 560M 700M SE +/- 467034.80, N = 3 SE +/- 1783157.89, N = 3 SE +/- 1202791.77, N = 3 SE +/- 794770.01, N = 3 SE +/- 1096202.13, N = 3 587219538.2 516770131.2 619106132.5 484098913.8 656484783.7 1. (CC) gcc options: -O3 -lm -ftree-vectorize -funroll-loops -lglut -lOpenCL -lGL
ViennaCL ViennaCL is an open-source linear algebra library written in C++ and with support for OpenCL and OpenMP. This test profile makes use of ViennaCL's built-in benchmarks. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org GB/s, More Is Better ViennaCL 1.7.1 Test: CPU BLAS - sCOPY NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 30 60 90 120 150 SE +/- 1.20, N = 3 SE +/- 1.20, N = 3 SE +/- 0.88, N = 3 SE +/- 1.20, N = 3 SE +/- 0.67, N = 3 132 131 132 132 107 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GB/s, More Is Better ViennaCL 1.7.1 Test: CPU BLAS - sAXPY NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 30 60 90 120 150 SE +/- 2.19, N = 3 SE +/- 4.81, N = 3 SE +/- 2.00, N = 3 SE +/- 0.33, N = 3 SE +/- 0.33, N = 3 156 153 156 154 120 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GB/s, More Is Better ViennaCL 1.7.1 Test: CPU BLAS - sDOT NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 40 80 120 160 200 SE +/- 2.73, N = 3 SE +/- 3.76, N = 3 SE +/- 2.40, N = 3 SE +/- 35.40, N = 3 SE +/- 0.58, N = 3 165.0 166.0 168.0 132.1 129.0 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GB/s, More Is Better ViennaCL 1.7.1 Test: CPU BLAS - dCOPY NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 16 32 48 64 80 SE +/- 0.32, N = 3 SE +/- 0.25, N = 3 SE +/- 0.74, N = 3 SE +/- 0.72, N = 3 SE +/- 0.18, N = 3 70.8 71.0 71.3 70.2 52.7 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GB/s, More Is Better ViennaCL 1.7.1 Test: CPU BLAS - dAXPY NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 20 40 60 80 100 SE +/- 0.12, N = 3 SE +/- 0.44, N = 3 SE +/- 0.57, N = 3 SE +/- 0.94, N = 3 SE +/- 0.12, N = 3 87.2 86.8 87.3 86.2 64.3 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GB/s, More Is Better ViennaCL 1.7.1 Test: CPU BLAS - dDOT NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 20 40 60 80 100 SE +/- 0.09, N = 3 SE +/- 0.22, N = 3 SE +/- 0.58, N = 3 SE +/- 0.84, N = 3 SE +/- 0.19, N = 3 96.8 96.7 96.4 95.2 70.8 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GB/s, More Is Better ViennaCL 1.7.1 Test: CPU BLAS - dGEMV-N NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 20 40 60 80 100 SE +/- 0.33, N = 3 SE +/- 0.33, N = 3 SE +/- 0.33, N = 3 SE +/- 0.88, N = 3 SE +/- 0.46, N = 3 102.0 103.0 103.0 103.0 78.5 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GB/s, More Is Better ViennaCL 1.7.1 Test: CPU BLAS - dGEMV-T NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 20 40 60 80 100 SE +/- 0.33, N = 3 SE +/- 0.33, N = 3 SE +/- 6.30, N = 3 SE +/- 0.33, N = 3 SE +/- 0.47, N = 3 109.0 109.0 102.7 110.0 82.6 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GFLOPs/s, More Is Better ViennaCL 1.7.1 Test: CPU BLAS - dGEMM-NN NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 30 60 90 120 150 SE +/- 4.04, N = 3 SE +/- 1.86, N = 3 SE +/- 1.15, N = 3 SE +/- 1.86, N = 3 SE +/- 1.50, N = 2 119 122 117 113 122 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GFLOPs/s, More Is Better ViennaCL 1.7.1 Test: CPU BLAS - dGEMM-NT NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 30 60 90 120 150 SE +/- 2.08, N = 3 SE +/- 1.76, N = 3 SE +/- 1.20, N = 3 SE +/- 3.28, N = 3 SE +/- 3.50, N = 2 117 122 118 119 119 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GFLOPs/s, More Is Better ViennaCL 1.7.1 Test: CPU BLAS - dGEMM-TN NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 30 60 90 120 150 SE +/- 1.00, N = 2 SE +/- 2.31, N = 3 SE +/- 2.08, N = 3 SE +/- 2.08, N = 3 SE +/- 3.00, N = 2 115 121 125 121 120 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GFLOPs/s, More Is Better ViennaCL 1.7.1 Test: CPU BLAS - dGEMM-TT NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 30 60 90 120 150 SE +/- 2.08, N = 3 SE +/- 1.20, N = 3 SE +/- 2.08, N = 3 SE +/- 0.88, N = 3 SE +/- 2.91, N = 3 122 118 124 113 117 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GB/s, More Is Better ViennaCL 1.7.1 Test: OpenCL BLAS - sCOPY NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 80 160 240 320 400 SE +/- 0.33, N = 3 SE +/- 0.33, N = 3 SE +/- 0.00, N = 3 SE +/- 1.00, N = 3 SE +/- 0.33, N = 3 334 330 336 363 373 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GB/s, More Is Better ViennaCL 1.7.1 Test: OpenCL BLAS - sAXPY NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 110 220 330 440 550 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.58, N = 3 SE +/- 1.20, N = 3 392 389 393 498 469 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GB/s, More Is Better ViennaCL 1.7.1 Test: OpenCL BLAS - sDOT NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 90 180 270 360 450 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.58, N = 3 SE +/- 1.00, N = 3 370 362 365 376 410 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GB/s, More Is Better ViennaCL 1.7.1 Test: OpenCL BLAS - dCOPY NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 130 260 390 520 650 SE +/- 0.33, N = 3 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.58, N = 3 SE +/- 4.00, N = 3 423 423 424 605 512 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GB/s, More Is Better ViennaCL 1.7.1 Test: OpenCL BLAS - dAXPY NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 160 320 480 640 800 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.33, N = 3 SE +/- 0.58, N = 3 SE +/- 0.00, N = 3 437 455 437 724 585 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GB/s, More Is Better ViennaCL 1.7.1 Test: OpenCL BLAS - dDOT NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 140 280 420 560 700 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.88, N = 3 SE +/- 1.33, N = 3 458 456 457 659 575 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GB/s, More Is Better ViennaCL 1.7.1 Test: OpenCL BLAS - dGEMV-N NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 50 100 150 200 250 SE +/- 0.33, N = 3 SE +/- 0.33, N = 3 SE +/- 0.33, N = 3 SE +/- 0.33, N = 3 SE +/- 0.00, N = 3 210 209 211 187 218 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GB/s, More Is Better ViennaCL 1.7.1 Test: OpenCL BLAS - dGEMV-T NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 90 180 270 360 450 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.33, N = 3 SE +/- 0.33, N = 3 389 387 391 374 424 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GFLOPs/s, More Is Better ViennaCL 1.7.1 Test: OpenCL BLAS - dGEMM-NN NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 150 300 450 600 750 SE +/- 0.00, N = 3 SE +/- 0.33, N = 3 SE +/- 0.33, N = 3 SE +/- 2.31, N = 3 SE +/- 1.33, N = 3 577 473 604 592 681 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GFLOPs/s, More Is Better ViennaCL 1.7.1 Test: OpenCL BLAS - dGEMM-NT NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 150 300 450 600 750 SE +/- 0.00, N = 3 SE +/- 0.33, N = 3 SE +/- 0.33, N = 3 SE +/- 2.33, N = 3 SE +/- 1.00, N = 3 584 477 612 595 689 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GFLOPs/s, More Is Better ViennaCL 1.7.1 Test: OpenCL BLAS - dGEMM-TN NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 150 300 450 600 750 SE +/- 0.00, N = 3 SE +/- 0.33, N = 3 SE +/- 0.67, N = 3 SE +/- 2.03, N = 3 SE +/- 1.00, N = 3 599 494 634 594 714 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
OpenBenchmarking.org GFLOPs/s, More Is Better ViennaCL 1.7.1 Test: OpenCL BLAS - dGEMM-TT NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 160 320 480 640 800 SE +/- 0.00, N = 3 SE +/- 0.33, N = 3 SE +/- 0.33, N = 3 SE +/- 1.33, N = 3 613 502 648 593 731 1. (CXX) g++ options: -fopenmp -O3 -rdynamic -lOpenCL
Libplacebo Libplacebo is a multimedia rendering library based on the core rendering code of the MPV player. The libplacebo benchmark relies on the Vulkan API and tests various primitives. Learn more via the OpenBenchmarking.org test page.
NVIDIA RTX 4070 SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status. E: libplacebo: line 3: ./src/bench: No such file or directory
OpenBenchmarking.org FPS, More Is Better Libplacebo 5.229.1 Test: deband_heavy NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER 500 1000 1500 2000 2500 SE +/- 0.15, N = 3 SE +/- 0.40, N = 3 SE +/- 2.90, N = 3 SE +/- 0.75, N = 3 SE +/- 2.26, N = 3 1844.08 2306.56 2015.93 2495.92 2186.70 1. (CXX) g++ options: -lm -pthread -ldl -fvisibility=hidden -std=c++20 -O2 -fno-math-errno -fPIC -MD -MQ -MF
OpenBenchmarking.org FPS, More Is Better Libplacebo 5.229.1 Test: polar_nocompute NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER 600 1200 1800 2400 3000 SE +/- 0.29, N = 3 SE +/- 1.70, N = 3 SE +/- 3.16, N = 3 SE +/- 1.94, N = 3 SE +/- 0.24, N = 3 1969.19 2459.03 2116.50 2653.03 2327.55 1. (CXX) g++ options: -lm -pthread -ldl -fvisibility=hidden -std=c++20 -O2 -fno-math-errno -fPIC -MD -MQ -MF
OpenBenchmarking.org FPS, More Is Better Libplacebo 5.229.1 Test: hdr_peakdetect NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER 1100 2200 3300 4400 5500 SE +/- 144.09, N = 3 SE +/- 165.74, N = 3 SE +/- 13.97, N = 3 SE +/- 13.83, N = 3 SE +/- 3.65, N = 3 3452.43 3475.06 5104.10 3913.34 3292.37 1. (CXX) g++ options: -lm -pthread -ldl -fvisibility=hidden -std=c++20 -O2 -fno-math-errno -fPIC -MD -MQ -MF
OpenBenchmarking.org FPS, More Is Better Libplacebo 5.229.1 Test: hdr_lut NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER 900 1800 2700 3600 4500 SE +/- 0.24, N = 3 SE +/- 33.96, N = 3 SE +/- 22.23, N = 3 SE +/- 16.37, N = 3 SE +/- 12.09, N = 3 3940.40 3976.04 3376.85 3822.16 3905.98 1. (CXX) g++ options: -lm -pthread -ldl -fvisibility=hidden -std=c++20 -O2 -fno-math-errno -fPIC -MD -MQ -MF
OpenBenchmarking.org FPS, More Is Better Libplacebo 5.229.1 Test: av1_grain_lap NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER NVIDIA RTX 4070 SUPER 900 1800 2700 3600 4500 SE +/- 72.31, N = 3 SE +/- 35.33, N = 3 SE +/- 42.82, N = 3 SE +/- 48.74, N = 3 SE +/- 5.52, N = 3 4126.40 4143.96 4096.48 4044.72 4171.00 1. (CXX) g++ options: -lm -pthread -ldl -fvisibility=hidden -std=c++20 -O2 -fno-math-errno -fPIC -MD -MQ -MF
RealSR-NCNN RealSR-NCNN is an NCNN neural network implementation of the RealSR project and accelerated using the Vulkan API. RealSR is the Real-World Super Resolution via Kernel Estimation and Noise Injection. NCNN is a high performance neural network inference framework optimized for mobile and other platforms developed by Tencent. This test profile times how long it takes to increase the resolution of a sample image by a scale of 4x with Vulkan. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org Seconds, Fewer Is Better RealSR-NCNN 20200818 Scale: 4x - TAA: No NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 2 4 6 8 10 SE +/- 0.150, N = 15 SE +/- 0.006, N = 3 SE +/- 0.039, N = 3 SE +/- 0.016, N = 3 SE +/- 0.003, N = 3 6.323 7.092 5.962 5.556 5.633
OpenBenchmarking.org Seconds, Fewer Is Better RealSR-NCNN 20200818 Scale: 4x - TAA: Yes NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 10 20 30 40 50 SE +/- 0.02, N = 3 SE +/- 0.23, N = 3 SE +/- 0.03, N = 3 SE +/- 0.06, N = 3 SE +/- 0.02, N = 3 34.89 42.85 33.63 30.31 30.72
VkFFT OpenBenchmarking.org Benchmark Score, More Is Better VkFFT 1.2.31 Test: FFT + iFFT R2C / C2R NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 13K 26K 39K 52K 65K SE +/- 702.53, N = 15 SE +/- 745.02, N = 13 SE +/- 520.37, N = 3 SE +/- 320.62, N = 3 SE +/- 772.47, N = 15 54794 47097 55446 48418 59378 1. (CXX) g++ options: -O3 -lrt
OpenBenchmarking.org Benchmark Score, More Is Better VkFFT 1.2.31 Test: FFT + iFFT C2C 1D batched in half precision NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 60K 120K 180K 240K 300K SE +/- 159.17, N = 3 SE +/- 1301.92, N = 3 SE +/- 1708.38, N = 3 SE +/- 160.60, N = 3 SE +/- 3524.05, N = 12 131705 137762 136210 273221 143992 1. (CXX) g++ options: -O3 -lrt
OpenBenchmarking.org Benchmark Score, More Is Better VkFFT 1.2.31 Test: FFT + iFFT C2C Bluestein in single precision NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 3K 6K 9K 12K 15K SE +/- 102.52, N = 3 SE +/- 52.09, N = 3 SE +/- 118.41, N = 3 SE +/- 115.62, N = 3 SE +/- 73.00, N = 3 15166 13714 15125 14205 16141 1. (CXX) g++ options: -O3 -lrt
OpenBenchmarking.org Benchmark Score, More Is Better VkFFT 1.2.31 Test: FFT + iFFT C2C 1D batched in double precision NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 7K 14K 21K 28K 35K SE +/- 146.69, N = 3 SE +/- 125.94, N = 3 SE +/- 302.46, N = 3 SE +/- 50.66, N = 3 SE +/- 325.03, N = 3 24317 22390 25431 30912 27947 1. (CXX) g++ options: -O3 -lrt
OpenBenchmarking.org Benchmark Score, More Is Better VkFFT 1.2.31 Test: FFT + iFFT C2C 1D batched in single precision NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 30K 60K 90K 120K 150K SE +/- 7.94, N = 3 SE +/- 13.72, N = 3 SE +/- 0.88, N = 3 SE +/- 9.64, N = 3 SE +/- 33.60, N = 3 73929 77774 73942 141876 104003 1. (CXX) g++ options: -O3 -lrt
OpenBenchmarking.org Benchmark Score, More Is Better VkFFT 1.2.31 Test: FFT + iFFT C2C multidimensional in single precision NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 13K 26K 39K 52K 65K SE +/- 407.19, N = 15 SE +/- 476.57, N = 5 SE +/- 417.77, N = 15 SE +/- 407.28, N = 15 SE +/- 251.10, N = 3 50299 47212 51528 50856 59790 1. (CXX) g++ options: -O3 -lrt
OpenBenchmarking.org Benchmark Score, More Is Better VkFFT 1.2.31 Test: FFT + iFFT C2C Bluestein benchmark in double precision NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 1100 2200 3300 4400 5500 SE +/- 12.55, N = 3 SE +/- 4.51, N = 3 SE +/- 11.35, N = 3 SE +/- 9.84, N = 3 SE +/- 11.37, N = 3 4451 3886 4647 4195 5047 1. (CXX) g++ options: -O3 -lrt
OpenBenchmarking.org Benchmark Score, More Is Better VkFFT 1.2.31 Test: FFT + iFFT C2C 1D batched in single precision, no reshuffling NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 30K 60K 90K 120K 150K SE +/- 37.77, N = 3 SE +/- 5.84, N = 3 SE +/- 28.54, N = 3 SE +/- 37.44, N = 3 SE +/- 20.80, N = 3 75078 79057 75141 144311 105549 1. (CXX) g++ options: -O3 -lrt
vkpeak Vkpeak is a Vulkan compute benchmark inspired by OpenCL's clpeak. Vkpeak provides Vulkan compute performance measurements for FP16 / FP32 / FP64 / INT16 / INT32 scalar and vec4 performance. Learn more via the OpenBenchmarking.org test page.
NVIDIA RTX 4070 SUPER: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status.
NVIDIA RTX 4070: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status.
NVIDIA RTX 4070 TI: The test quit with a non-zero exit status. The test quit with a non-zero exit status. The test quit with a non-zero exit status.
ProjectPhysX OpenCL-Benchmark ProjectPhysX OpenCL-Benchmark provides various OpenCL compute and memory bandwidth micro-benchmarks Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org TFLOPs/s, More Is Better ProjectPhysX OpenCL-Benchmark 1.2 Operation: FP64 Compute NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 0.1672 0.3344 0.5016 0.6688 0.836 SE +/- 0.000, N = 3 SE +/- 0.001, N = 3 SE +/- 0.001, N = 3 SE +/- 0.001, N = 3 SE +/- 0.001, N = 3 0.621 0.510 0.660 0.637 0.743 1. (CXX) g++ options: -std=c++17 -pthread -lOpenCL
VkResample VkResample is a Vulkan-based image upscaling library based on VkFFT. The sample input file is upscaling a 4K image to 8K using Vulkan-based GPU acceleration. Learn more via the OpenBenchmarking.org test page.
OpenBenchmarking.org ms, Fewer Is Better VkResample 1.0 Upscale: 2x - Precision: Double NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 90 180 270 360 450 SE +/- 0.30, N = 3 SE +/- 0.77, N = 3 SE +/- 0.35, N = 3 SE +/- 0.30, N = 3 SE +/- 0.02, N = 3 339.59 415.16 322.06 333.64 285.99 1. (CXX) g++ options: -O3
OpenBenchmarking.org ms, Fewer Is Better VkResample 1.0 Upscale: 2x - Precision: Single NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 5 10 15 20 25 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 18.49 18.02 18.46 10.32 13.36 1. (CXX) g++ options: -O3
Waifu2x-NCNN Vulkan Waifu2x-NCNN is an NCNN neural network implementation of the Waifu2x converter project and accelerated using the Vulkan API. NCNN is a high performance neural network inference framework optimized for mobile and other platforms developed by Tencent. This test profile times how long it takes to increase the resolution of a sample image with Vulkan. Learn more via the OpenBenchmarking.org test page.
Scale: 2x - Denoise: 3 - TAA: No
NVIDIA RTX 4070 SUPER: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 4070: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 4070 TI: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 3090: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
NVIDIA RTX 4070 TI SUPER: The test run did not produce a result. The test run did not produce a result. The test run did not produce a result.
OpenBenchmarking.org Seconds, Fewer Is Better Waifu2x-NCNN Vulkan 20200818 Scale: 2x - Denoise: 3 - TAA: Yes NVIDIA RTX 4070 SUPER NVIDIA RTX 4070 NVIDIA RTX 4070 TI NVIDIA RTX 3090 NVIDIA RTX 4070 TI SUPER 0.7205 1.441 2.1615 2.882 3.6025 SE +/- 0.014, N = 3 SE +/- 0.028, N = 3 SE +/- 0.009, N = 3 SE +/- 0.011, N = 3 SE +/- 0.028, N = 3 2.855 3.168 2.854 3.202 2.660
NVIDIA RTX 4070 SUPER Processor: Intel Core i9-13900K @ 5.50GHz (24 Cores / 32 Threads), Motherboard: ASUS TUF GAMING Z790-PRO WIFI (1401 BIOS), Chipset: Intel Device 7a27, Memory: 32GB, Disk: 4001GB Seagate ZP4000GP304001, Graphics: ASUS NVIDIA GeForce RTX 4070 SUPER 12GB, Audio: Realtek ALC1220, Monitor: ARZOPA, Network: Intel I226-V + Intel Device 7a70
OS: EndeavourOS rolling, Kernel: 6.7.1-arch1-1 (x86_64), Desktop: KDE Plasma 5.27.10, Display Server: X Server 1.21.1.11, Display Driver: NVIDIA 550.40.07, OpenGL: 4.6.0, OpenCL: OpenCL 3.0 CUDA 12.4.74, Compiler: GCC 13.2.1 20230801, File-System: ext4, Screen Resolution: 1920x1080
Kernel Notes: Transparent Huge Pages: alwaysCompiler Notes: --disable-libssp --disable-libstdcxx-pch --disable-werror --enable-__cxa_atexit --enable-bootstrap --enable-cet=auto --enable-checking=release --enable-clocale=gnu --enable-default-pie --enable-default-ssp --enable-gnu-indirect-function --enable-gnu-unique-object --enable-languages=ada,c,c++,d,fortran,go,lto,objc,obj-c++ --enable-libstdcxx-backtrace --enable-link-serialization=1 --enable-lto --enable-multilib --enable-plugin --enable-shared --enable-threads=posix --mandir=/usr/share/man --with-build-config=bootstrap-lto --with-linker-hash-style=gnuProcessor Notes: Scaling Governor: intel_pstate powersave (EPP: balance_performance) - CPU Microcode: 0x11dGraphics Notes: BAR1 / Visible vRAM Size: 16384 MiB - vBIOS Version: 95.04.69.00.c1Security Notes: gather_data_sampling: Not affected + itlb_multihit: Not affected + l1tf: Not affected + mds: Not affected + meltdown: Not affected + mmio_stale_data: Not affected + retbleed: Not affected + spec_rstack_overflow: Not affected + spec_store_bypass: Mitigation of SSB disabled via prctl + spectre_v1: Mitigation of usercopy/swapgs barriers and __user pointer sanitization + spectre_v2: Mitigation of Enhanced / Automatic IBRS IBPB: conditional RSB filling PBRSB-eIBRS: SW sequence + srbds: Not affected + tsx_async_abort: Not affected
Testing initiated at 25 January 2024 21:36 by user test.
NVIDIA RTX 4070 Processor: Intel Core i9-13900K @ 5.50GHz (24 Cores / 32 Threads), Motherboard: ASUS TUF GAMING Z790-PRO WIFI (1401 BIOS), Chipset: Intel Device 7a27, Memory: 32GB, Disk: 4001GB Seagate ZP4000GP304001, Graphics: MSI NVIDIA GeForce RTX 4070 12GB, Audio: Realtek ALC1220, Monitor: ARZOPA, Network: Intel I226-V + Intel Device 7a70
OS: EndeavourOS rolling, Kernel: 6.7.1-arch1-1 (x86_64), Desktop: KDE Plasma 5.27.10, Display Server: X Server 1.21.1.11, Display Driver: NVIDIA 550.40.07, OpenGL: 4.6.0, OpenCL: OpenCL 3.0 CUDA 12.4.74, Compiler: GCC 13.2.1 20230801 + CUDA 12.3, File-System: ext4, Screen Resolution: 1920x1080
Kernel Notes: Transparent Huge Pages: alwaysEnvironment Notes: NVCC_PREPEND_FLAGS="-ccbin /opt/cuda/bin"Compiler Notes: --disable-libssp --disable-libstdcxx-pch --disable-werror --enable-__cxa_atexit --enable-bootstrap --enable-cet=auto --enable-checking=release --enable-clocale=gnu --enable-default-pie --enable-default-ssp --enable-gnu-indirect-function --enable-gnu-unique-object --enable-languages=ada,c,c++,d,fortran,go,lto,objc,obj-c++ --enable-libstdcxx-backtrace --enable-link-serialization=1 --enable-lto --enable-multilib --enable-plugin --enable-shared --enable-threads=posix --mandir=/usr/share/man --with-build-config=bootstrap-lto --with-linker-hash-style=gnuProcessor Notes: Scaling Governor: intel_pstate powersave (EPP: balance_performance) - CPU Microcode: 0x11dGraphics Notes: BAR1 / Visible vRAM Size: 16384 MiB - vBIOS Version: 95.04.3e.40.2aPython Notes: Python 3.11.6Security Notes: gather_data_sampling: Not affected + itlb_multihit: Not affected + l1tf: Not affected + mds: Not affected + meltdown: Not affected + mmio_stale_data: Not affected + retbleed: Not affected + spec_rstack_overflow: Not affected + spec_store_bypass: Mitigation of SSB disabled via prctl + spectre_v1: Mitigation of usercopy/swapgs barriers and __user pointer sanitization + spectre_v2: Mitigation of Enhanced / Automatic IBRS IBPB: conditional RSB filling PBRSB-eIBRS: SW sequence + srbds: Not affected + tsx_async_abort: Not affected
Testing initiated at 28 January 2024 13:02 by user test.
NVIDIA RTX 4070 TI Processor: Intel Core i9-13900K @ 5.50GHz (24 Cores / 32 Threads), Motherboard: ASUS TUF GAMING Z790-PRO WIFI (1401 BIOS), Chipset: Intel Device 7a27, Memory: 32GB, Disk: 4001GB Seagate ZP4000GP304001, Graphics: NVIDIA GeForce RTX 4070 Ti 12GB, Audio: Realtek ALC1220, Monitor: ARZOPA, Network: Intel I226-V + Intel Device 7a70
OS: EndeavourOS rolling, Kernel: 6.7.1-arch1-1 (x86_64), Desktop: KDE Plasma 5.27.10, Display Server: X Server 1.21.1.11, Display Driver: NVIDIA 550.40.07, OpenGL: 4.6.0, OpenCL: OpenCL 3.0 CUDA 12.4.74, Compiler: GCC 13.2.1 20230801 + CUDA 12.3, File-System: ext4, Screen Resolution: 1920x1080
Kernel Notes: Transparent Huge Pages: alwaysEnvironment Notes: NVCC_PREPEND_FLAGS="-ccbin /opt/cuda/bin"Compiler Notes: --disable-libssp --disable-libstdcxx-pch --disable-werror --enable-__cxa_atexit --enable-bootstrap --enable-cet=auto --enable-checking=release --enable-clocale=gnu --enable-default-pie --enable-default-ssp --enable-gnu-indirect-function --enable-gnu-unique-object --enable-languages=ada,c,c++,d,fortran,go,lto,objc,obj-c++ --enable-libstdcxx-backtrace --enable-link-serialization=1 --enable-lto --enable-multilib --enable-plugin --enable-shared --enable-threads=posix --mandir=/usr/share/man --with-build-config=bootstrap-lto --with-linker-hash-style=gnuProcessor Notes: Scaling Governor: intel_pstate powersave (EPP: balance_performance) - CPU Microcode: 0x11dGraphics Notes: BAR1 / Visible vRAM Size: 16384 MiB - vBIOS Version: 95.04.31.00.36Python Notes: Python 3.11.6Security Notes: gather_data_sampling: Not affected + itlb_multihit: Not affected + l1tf: Not affected + mds: Not affected + meltdown: Not affected + mmio_stale_data: Not affected + retbleed: Not affected + spec_rstack_overflow: Not affected + spec_store_bypass: Mitigation of SSB disabled via prctl + spectre_v1: Mitigation of usercopy/swapgs barriers and __user pointer sanitization + spectre_v2: Mitigation of Enhanced / Automatic IBRS IBPB: conditional RSB filling PBRSB-eIBRS: SW sequence + srbds: Not affected + tsx_async_abort: Not affected
Testing initiated at 29 January 2024 17:08 by user test.
NVIDIA RTX 3090 Processor: Intel Core i9-13900K @ 5.50GHz (24 Cores / 32 Threads), Motherboard: ASUS TUF GAMING Z790-PRO WIFI (1401 BIOS), Chipset: Intel Device 7a27, Memory: 32GB, Disk: 4001GB Seagate ZP4000GP304001, Graphics: NVIDIA GeForce RTX 3090 24GB, Audio: Realtek ALC1220, Monitor: PI-KVM Video, Network: Intel I226-V + Intel Device 7a70
OS: EndeavourOS rolling, Kernel: 6.7.4-arch1-1 (x86_64), Desktop: KDE Plasma 5.27.10, Display Server: X Server 1.21.1.11, Display Driver: NVIDIA 550.40.07, OpenGL: 4.6.0, OpenCL: OpenCL 3.0 CUDA 12.4.74, Compiler: GCC 13.2.1 20230801 + CUDA 12.3, File-System: ext4, Screen Resolution: 1920x1080
Kernel Notes: Transparent Huge Pages: alwaysEnvironment Notes: NVCC_PREPEND_FLAGS="-ccbin /opt/cuda/bin"Compiler Notes: --disable-libssp --disable-libstdcxx-pch --disable-werror --enable-__cxa_atexit --enable-bootstrap --enable-cet=auto --enable-checking=release --enable-clocale=gnu --enable-default-pie --enable-default-ssp --enable-gnu-indirect-function --enable-gnu-unique-object --enable-languages=ada,c,c++,d,fortran,go,lto,m2,objc,obj-c++ --enable-libstdcxx-backtrace --enable-link-serialization=1 --enable-lto --enable-multilib --enable-plugin --enable-shared --enable-threads=posix --mandir=/usr/share/man --with-build-config=bootstrap-lto --with-linker-hash-style=gnuProcessor Notes: Scaling Governor: intel_pstate powersave (EPP: balance_performance) - CPU Microcode: 0x11dGraphics Notes: BAR1 / Visible vRAM Size: 256 MiB - vBIOS Version: 94.02.26.08.baPython Notes: Python 3.11.6Security Notes: gather_data_sampling: Not affected + itlb_multihit: Not affected + l1tf: Not affected + mds: Not affected + meltdown: Not affected + mmio_stale_data: Not affected + retbleed: Not affected + spec_rstack_overflow: Not affected + spec_store_bypass: Mitigation of SSB disabled via prctl + spectre_v1: Mitigation of usercopy/swapgs barriers and __user pointer sanitization + spectre_v2: Mitigation of Enhanced / Automatic IBRS IBPB: conditional RSB filling PBRSB-eIBRS: SW sequence + srbds: Not affected + tsx_async_abort: Not affected
Testing initiated at 7 February 2024 20:29 by user saddytech.
NVIDIA RTX 4070 TI SUPER Processor: Intel Core i9-13900K @ 5.50GHz (24 Cores / 32 Threads), Motherboard: ASUS TUF GAMING Z790-PRO WIFI (1630 BIOS), Chipset: Intel Raptor Lake-S PCH, Memory: 32GB, Disk: 4001GB Seagate ZP4000GP304001 + 0GB CD-ROM Drive, Graphics: ASUS NVIDIA GeForce RTX 4070 Ti SUPER 16GB, Audio: Realtek ALC1220, Monitor: PI-KVM Video, Network: Intel I226-V + Intel Raptor Lake-S PCH CNVi WiFi
OS: EndeavourOS rolling, Kernel: 6.7.4-arch1-1 (x86_64), Desktop: KDE Plasma 5.27.10, Display Server: X Server 1.21.1.11, Display Driver: NVIDIA 550.40.07, OpenGL: 4.6.0, OpenCL: OpenCL 2.1 AMD-APP (3602.0) + OpenCL 3.0 CUDA 12.4.74, Compiler: GCC 13.2.1 20230801 + CUDA 12.3, File-System: ext4, Screen Resolution: 1920x1080
Kernel Notes: Transparent Huge Pages: alwaysEnvironment Notes: NVCC_PREPEND_FLAGS="-ccbin /opt/cuda/bin"Compiler Notes: --disable-libssp --disable-libstdcxx-pch --disable-werror --enable-__cxa_atexit --enable-bootstrap --enable-cet=auto --enable-checking=release --enable-clocale=gnu --enable-default-pie --enable-default-ssp --enable-gnu-indirect-function --enable-gnu-unique-object --enable-languages=ada,c,c++,d,fortran,go,lto,m2,objc,obj-c++ --enable-libstdcxx-backtrace --enable-link-serialization=1 --enable-lto --enable-multilib --enable-plugin --enable-shared --enable-threads=posix --mandir=/usr/share/man --with-build-config=bootstrap-lto --with-linker-hash-style=gnuProcessor Notes: Scaling Governor: intel_pstate powersave (EPP: balance_performance) - CPU Microcode: 0x11fGraphics Notes: BAR1 / Visible vRAM Size: 256 MiB - vBIOS Version: 95.03.45.00.c5Python Notes: Python 3.11.7Security Notes: gather_data_sampling: Not affected + itlb_multihit: Not affected + l1tf: Not affected + mds: Not affected + meltdown: Not affected + mmio_stale_data: Not affected + retbleed: Not affected + spec_rstack_overflow: Not affected + spec_store_bypass: Mitigation of SSB disabled via prctl + spectre_v1: Mitigation of usercopy/swapgs barriers and __user pointer sanitization + spectre_v2: Mitigation of Enhanced / Automatic IBRS IBPB: conditional RSB filling PBRSB-eIBRS: SW sequence + srbds: Not affected + tsx_async_abort: Not affected
Testing initiated at 15 February 2024 17:25 by user saddytech.