GCC 4.9 Linux LTO Optimizations Some link-time optimization compiler benchmarks by Michael Larabel for Phoronix.com of GCC 4.8.2 and 4.9.0 RC1. The Link-time optimization results with -flto didn't turn out to be as exciting as anticipated, so here they are for this short future article on phoronix. Benchmarks from an Intel Core i7 Haswell running Ubuntu Linux.
HTML result view exported from: https://openbenchmarking.org/result/1404126-PTS-GCC4849L62&sor&gru .
GCC 4.9 Linux LTO Optimizations Processor Motherboard Chipset Memory Disk Graphics Audio Monitor Network OS Kernel Desktop Display Server Display Driver OpenGL Compiler File-System Screen Resolution GCC 4.8.2 - Stock GCC 4.8.2 - LTO GCC 4.9.0 RC1 - Stock GCC 4.9.0 RC1 - LTO Intel Core i7-4770K @ 3.50GHz (8 Cores) ECS Z87H3-A2X EXTREME v1.0 Intel 4th Gen Core DRAM 16384MB 120GB Samsung SSD 840 ECS NVIDIA GeForce GTX 460 768MB (675/1804MHz) Realtek ALC1150 Samsung SyncMaster Realtek RTL8111/8168/8411 Ubuntu 14.04 3.13.0-22-generic (x86_64) Unity 7.2.0 X Server 1.15.0 NVIDIA 337.12 4.3.0 GCC 4.8.2 ext4 2560x1600 GCC 4.9.0 20140411 OpenBenchmarking.org Compiler Details - --enable-checking=release Processor Details - Scaling Governor: acpi-cpufreq ondemand
GCC 4.9 Linux LTO Optimizations hpcc: G-Ptrans hpcc: EP-STREAM Triad hpcc: G-HPL hpcc: G-Ffte hpcc: EP-DGEMM hpcc: G-Rand Access graphics-magick: Blur graphics-magick: Sharpen graphics-magick: Resizing graphics-magick: HWB Color Space graphics-magick: Local Adaptive Thresholding byte: Dhrystone 2 himeno: Poisson Pressure Solver hint: FLOAT apache: Static Web Page Serving ebizzy: Records/s build-apache: Time To Compile build-imagemagick: Time To Compile build-php: Time To Compile c-ray: Total Time open-porous-media: Upscale-Relperm smallpt: Global Illumination Renderer; 100 Samples encode-flac: WAV To FLAC encode-mp3: WAV To MP3 ffmpeg: H.264 HD To NTSC DV nero2d: Total Time GCC 4.8.2 - Stock GCC 4.8.2 - LTO GCC 4.9.0 RC1 - Stock GCC 4.9.0 RC1 - LTO 1.26591 1.08895 50.05477 2.11922 6.57694 0.00979 170 140 198 216 97 35908267.37 1810.10 367476347.00 35904.53 42849 27.14 60.35 26.27 16.98 100.58 24 4.82 12.42 13.99 453.45 1.26956 1.08612 49.78783 2.12967 6.63749 0.00980 171 141 199 218 97 34717223.53 1813.55 367633749.83 35521.51 42698 44.37 124.51 93.29 16.98 100.42 24 13.24 13.80 458.86 1.26904 1.08175 50.08117 2.11338 6.72611 0.00980 166 140 199 214 102 36694838.47 1828.15 373674384.83 35953.73 42950 27.88 60.59 27.05 17.09 3.70 10.87 13.80 1.26850 1.08126 49.96230 2.13555 6.63834 0.00989 165 140 198 213 102 33544122.87 1825.41 371516290.66 35817.51 42565 45.74 129.91 95.42 17.08 3.72 10.88 13.75 OpenBenchmarking.org
HPC Challenge Test / Class: G-Ptrans OpenBenchmarking.org GB/s, More Is Better HPC Challenge 1.4.3 Test / Class: G-Ptrans GCC 4.8.2 - LTO GCC 4.9.0 RC1 - Stock GCC 4.9.0 RC1 - LTO GCC 4.8.2 - Stock 0.2857 0.5714 0.8571 1.1428 1.4285 SE +/- 0.00112, N = 3 SE +/- 0.00329, N = 3 SE +/- 0.00220, N = 3 SE +/- 0.00051, N = 3 1.26956 1.26904 1.26850 1.26591 -flto -flto 1. (CC) gcc options: -lblas -lm -pthread -lmpi -ldl -lhwloc -fomit-frame-pointer -O3 -march=native -funroll-loops 2. BLAS + Open MPI 1.6.5
HPC Challenge Test / Class: EP-STREAM Triad OpenBenchmarking.org GB/s, More Is Better HPC Challenge 1.4.3 Test / Class: EP-STREAM Triad GCC 4.8.2 - Stock GCC 4.8.2 - LTO GCC 4.9.0 RC1 - Stock GCC 4.9.0 RC1 - LTO 0.245 0.49 0.735 0.98 1.225 SE +/- 0.00488, N = 3 SE +/- 0.00405, N = 3 SE +/- 0.00379, N = 3 SE +/- 0.00342, N = 3 1.08895 1.08612 1.08175 1.08126 -flto -flto 1. (CC) gcc options: -lblas -lm -pthread -lmpi -ldl -lhwloc -fomit-frame-pointer -O3 -march=native -funroll-loops 2. BLAS + Open MPI 1.6.5
HPC Challenge Test / Class: G-HPL OpenBenchmarking.org GFLOPS, More Is Better HPC Challenge 1.4.3 Test / Class: G-HPL GCC 4.9.0 RC1 - Stock GCC 4.8.2 - Stock GCC 4.9.0 RC1 - LTO GCC 4.8.2 - LTO 11 22 33 44 55 SE +/- 0.05, N = 3 SE +/- 0.08, N = 3 SE +/- 0.11, N = 3 SE +/- 0.24, N = 3 50.08 50.05 49.96 49.79 -flto -flto 1. (CC) gcc options: -lblas -lm -pthread -lmpi -ldl -lhwloc -fomit-frame-pointer -O3 -march=native -funroll-loops 2. BLAS + Open MPI 1.6.5
HPC Challenge Test / Class: G-Ffte OpenBenchmarking.org GFLOPS, More Is Better HPC Challenge 1.4.3 Test / Class: G-Ffte GCC 4.9.0 RC1 - LTO GCC 4.8.2 - LTO GCC 4.8.2 - Stock GCC 4.9.0 RC1 - Stock 0.4805 0.961 1.4415 1.922 2.4025 SE +/- 0.01396, N = 3 SE +/- 0.00327, N = 3 SE +/- 0.00495, N = 3 SE +/- 0.00571, N = 3 2.13555 2.12967 2.11922 2.11338 -flto -flto 1. (CC) gcc options: -lblas -lm -pthread -lmpi -ldl -lhwloc -fomit-frame-pointer -O3 -march=native -funroll-loops 2. BLAS + Open MPI 1.6.5
HPC Challenge Test / Class: EP-DGEMM OpenBenchmarking.org GFLOPS, More Is Better HPC Challenge 1.4.3 Test / Class: EP-DGEMM GCC 4.9.0 RC1 - Stock GCC 4.9.0 RC1 - LTO GCC 4.8.2 - LTO GCC 4.8.2 - Stock 2 4 6 8 10 SE +/- 0.05569, N = 3 SE +/- 0.06945, N = 3 SE +/- 0.05029, N = 3 SE +/- 0.04345, N = 3 6.72611 6.63834 6.63749 6.57694 -flto -flto 1. (CC) gcc options: -lblas -lm -pthread -lmpi -ldl -lhwloc -fomit-frame-pointer -O3 -march=native -funroll-loops 2. BLAS + Open MPI 1.6.5
HPC Challenge Test / Class: G-Random Access OpenBenchmarking.org GUP/s, More Is Better HPC Challenge 1.4.3 Test / Class: G-Random Access GCC 4.9.0 RC1 - LTO GCC 4.9.0 RC1 - Stock GCC 4.8.2 - LTO GCC 4.8.2 - Stock 0.0022 0.0044 0.0066 0.0088 0.011 SE +/- 0.00003, N = 3 SE +/- 0.00013, N = 3 SE +/- 0.00001, N = 3 SE +/- 0.00008, N = 3 0.00989 0.00980 0.00980 0.00979 -flto -flto 1. (CC) gcc options: -lblas -lm -pthread -lmpi -ldl -lhwloc -fomit-frame-pointer -O3 -march=native -funroll-loops 2. BLAS + Open MPI 1.6.5
GraphicsMagick Operation: Blur OpenBenchmarking.org Iterations Per Minute, More Is Better GraphicsMagick 1.3.19 Operation: Blur GCC 4.8.2 - LTO GCC 4.8.2 - Stock GCC 4.9.0 RC1 - Stock GCC 4.9.0 RC1 - LTO 40 80 120 160 200 SE +/- 0.33, N = 3 SE +/- 0.33, N = 3 SE +/- 0.33, N = 3 SE +/- 0.33, N = 3 171 170 166 165 -flto -flto 1. (CC) gcc options: -std=gnu99 -fopenmp -O3 -march=native -pthread -ljbig -lwebp -ljpeg -lXext -lX11 -llzma -lxml2 -lz -lm -lpthread
GraphicsMagick Operation: Sharpen OpenBenchmarking.org Iterations Per Minute, More Is Better GraphicsMagick 1.3.19 Operation: Sharpen GCC 4.8.2 - LTO GCC 4.9.0 RC1 - LTO GCC 4.9.0 RC1 - Stock GCC 4.8.2 - Stock 30 60 90 120 150 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.33, N = 3 141 140 140 140 -flto -flto 1. (CC) gcc options: -std=gnu99 -fopenmp -O3 -march=native -pthread -ljbig -lwebp -ljpeg -lXext -lX11 -llzma -lxml2 -lz -lm -lpthread
GraphicsMagick Operation: Resizing OpenBenchmarking.org Iterations Per Minute, More Is Better GraphicsMagick 1.3.19 Operation: Resizing GCC 4.9.0 RC1 - Stock GCC 4.8.2 - LTO GCC 4.9.0 RC1 - LTO GCC 4.8.2 - Stock 40 80 120 160 200 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.33, N = 3 SE +/- 0.33, N = 3 199 199 198 198 -flto -flto 1. (CC) gcc options: -std=gnu99 -fopenmp -O3 -march=native -pthread -ljbig -lwebp -ljpeg -lXext -lX11 -llzma -lxml2 -lz -lm -lpthread
GraphicsMagick Operation: HWB Color Space OpenBenchmarking.org Iterations Per Minute, More Is Better GraphicsMagick 1.3.19 Operation: HWB Color Space GCC 4.8.2 - LTO GCC 4.8.2 - Stock GCC 4.9.0 RC1 - Stock GCC 4.9.0 RC1 - LTO 50 100 150 200 250 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.33, N = 3 SE +/- 0.00, N = 3 218 216 214 213 -flto -flto 1. (CC) gcc options: -std=gnu99 -fopenmp -O3 -march=native -pthread -ljbig -lwebp -ljpeg -lXext -lX11 -llzma -lxml2 -lz -lm -lpthread
GraphicsMagick Operation: Local Adaptive Thresholding OpenBenchmarking.org Iterations Per Minute, More Is Better GraphicsMagick 1.3.19 Operation: Local Adaptive Thresholding GCC 4.9.0 RC1 - LTO GCC 4.9.0 RC1 - Stock GCC 4.8.2 - LTO GCC 4.8.2 - Stock 20 40 60 80 100 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 SE +/- 0.33, N = 3 102 102 97 97 -flto -flto 1. (CC) gcc options: -std=gnu99 -fopenmp -O3 -march=native -pthread -ljbig -lwebp -ljpeg -lXext -lX11 -llzma -lxml2 -lz -lm -lpthread
BYTE Unix Benchmark Computational Test: Dhrystone 2 OpenBenchmarking.org LPS, More Is Better BYTE Unix Benchmark 3.6 Computational Test: Dhrystone 2 GCC 4.9.0 RC1 - Stock GCC 4.8.2 - Stock GCC 4.8.2 - LTO GCC 4.9.0 RC1 - LTO 8M 16M 24M 32M 40M SE +/- 165665.98, N = 3 SE +/- 13215.25, N = 3 SE +/- 9630.07, N = 3 SE +/- 32494.09, N = 3 36694838.47 35908267.37 34717223.53 33544122.87 -flto -flto 1. (CC) gcc options: -O3 -march=native
Himeno Benchmark Poisson Pressure Solver OpenBenchmarking.org MFLOPS, More Is Better Himeno Benchmark 3.0 Poisson Pressure Solver GCC 4.9.0 RC1 - Stock GCC 4.9.0 RC1 - LTO GCC 4.8.2 - LTO GCC 4.8.2 - Stock 400 800 1200 1600 2000 SE +/- 1.31, N = 3 SE +/- 0.63, N = 3 SE +/- 0.48, N = 3 SE +/- 3.93, N = 3 1828.15 1825.41 1813.55 1810.10 -flto -flto 1. (CC) gcc options: -O3 -march=native
Hierarchical INTegration Test: FLOAT OpenBenchmarking.org QUIPs, More Is Better Hierarchical INTegration 1.0 Test: FLOAT GCC 4.9.0 RC1 - Stock GCC 4.9.0 RC1 - LTO GCC 4.8.2 - LTO GCC 4.8.2 - Stock 80M 160M 240M 320M 400M SE +/- 362392.17, N = 3 SE +/- 994638.92, N = 3 SE +/- 129752.70, N = 3 SE +/- 103032.27, N = 3 373674384.83 371516290.66 367633749.83 367476347.00 -flto -flto 1. (CC) gcc options: -O3 -march=native -lm
Apache Benchmark Static Web Page Serving OpenBenchmarking.org Requests Per Second, More Is Better Apache Benchmark 2.4.7 Static Web Page Serving GCC 4.9.0 RC1 - Stock GCC 4.8.2 - Stock GCC 4.9.0 RC1 - LTO GCC 4.8.2 - LTO 8K 16K 24K 32K 40K SE +/- 76.61, N = 3 SE +/- 136.25, N = 3 SE +/- 55.60, N = 3 SE +/- 75.16, N = 3 35953.73 35904.53 35817.51 35521.51 -flto -flto 1. (CC) gcc options: -shared -fPIC -pthread -O3 -march=native
ebizzy Records/s OpenBenchmarking.org Seconds, More Is Better ebizzy 0.3 Records/s GCC 4.9.0 RC1 - Stock GCC 4.8.2 - Stock GCC 4.8.2 - LTO GCC 4.9.0 RC1 - LTO 9K 18K 27K 36K 45K SE +/- 44.08, N = 3 SE +/- 109.85, N = 3 SE +/- 104.13, N = 3 SE +/- 33.07, N = 3 42950 42849 42698 42565 -flto -flto 1. (CC) gcc options: -pthread -lpthread -O3 -march=native
Timed Apache Compilation Time To Compile OpenBenchmarking.org Seconds, Fewer Is Better Timed Apache Compilation 2.4.7 Time To Compile GCC 4.8.2 - Stock GCC 4.9.0 RC1 - Stock GCC 4.8.2 - LTO GCC 4.9.0 RC1 - LTO 10 20 30 40 50 SE +/- 0.19, N = 3 SE +/- 0.23, N = 3 SE +/- 0.07, N = 3 SE +/- 0.05, N = 3 27.14 27.88 44.37 45.74
Timed ImageMagick Compilation Time To Compile OpenBenchmarking.org Seconds, Fewer Is Better Timed ImageMagick Compilation 6.8.1-10 Time To Compile GCC 4.8.2 - Stock GCC 4.9.0 RC1 - Stock GCC 4.8.2 - LTO GCC 4.9.0 RC1 - LTO 30 60 90 120 150 SE +/- 0.36, N = 3 SE +/- 0.04, N = 3 SE +/- 0.30, N = 3 SE +/- 0.32, N = 3 60.35 60.59 124.51 129.91
Timed PHP Compilation Time To Compile OpenBenchmarking.org Seconds, Fewer Is Better Timed PHP Compilation 5.2.9 Time To Compile GCC 4.8.2 - Stock GCC 4.9.0 RC1 - Stock GCC 4.8.2 - LTO GCC 4.9.0 RC1 - LTO 20 40 60 80 100 SE +/- 0.08, N = 3 SE +/- 0.02, N = 3 SE +/- 1.16, N = 3 SE +/- 0.14, N = 3 26.27 27.05 93.29 95.42 -flto -flto 1. (CC) gcc options: -O3 -march=native -pedantic -ldl -lz -lm
C-Ray Total Time OpenBenchmarking.org Seconds, Fewer Is Better C-Ray 1.1 Total Time GCC 4.8.2 - Stock GCC 4.8.2 - LTO GCC 4.9.0 RC1 - LTO GCC 4.9.0 RC1 - Stock 4 8 12 16 20 SE +/- 0.01, N = 3 SE +/- 0.00, N = 3 SE +/- 0.01, N = 3 SE +/- 0.01, N = 3 16.98 16.98 17.08 17.09 -flto -flto 1. (CC) gcc options: -lm -lpthread -O3 -march=native
Open Porous Media OPM Benchmark: Upscale-Relperm OpenBenchmarking.org Seconds, Fewer Is Better Open Porous Media 2013-11-26 OPM Benchmark: Upscale-Relperm GCC 4.8.2 - LTO GCC 4.8.2 - Stock 20 40 60 80 100 SE +/- 0.38, N = 3 SE +/- 0.17, N = 3 100.42 100.58 1. (F9X) gfortran options: -rdynamic
Smallpt Global Illumination Renderer; 100 Samples OpenBenchmarking.org Seconds, Fewer Is Better Smallpt 1.0 Global Illumination Renderer; 100 Samples GCC 4.8.2 - Stock GCC 4.8.2 - LTO 6 12 18 24 30 SE +/- 0.00, N = 3 SE +/- 0.00, N = 3 24 24 -flto 1. (CXX) g++ options: -fopenmp -O3 -march=native
FLAC Audio Encoding WAV To FLAC OpenBenchmarking.org Seconds, Fewer Is Better FLAC Audio Encoding 1.3.0 WAV To FLAC GCC 4.9.0 RC1 - Stock GCC 4.9.0 RC1 - LTO GCC 4.8.2 - Stock 1.0845 2.169 3.2535 4.338 5.4225 SE +/- 0.09, N = 10 SE +/- 0.09, N = 10 SE +/- 0.06, N = 7 3.70 3.72 4.82 -flto 1. (CXX) g++ options: -O3 -march=native -fvisibility=hidden -logg -lm
LAME MP3 Encoding WAV To MP3 OpenBenchmarking.org Seconds, Fewer Is Better LAME MP3 Encoding 3.99.3 WAV To MP3 GCC 4.9.0 RC1 - Stock GCC 4.9.0 RC1 - LTO GCC 4.8.2 - Stock GCC 4.8.2 - LTO 3 6 9 12 15 SE +/- 0.01, N = 5 SE +/- 0.01, N = 5 SE +/- 0.04, N = 5 SE +/- 0.05, N = 5 10.87 10.88 12.42 13.24 -flto -flto 1. (CC) gcc options: -pipe -O3 -march=native -lm
FFmpeg H.264 HD To NTSC DV OpenBenchmarking.org Seconds, Fewer Is Better FFmpeg 2.1.1 H.264 HD To NTSC DV GCC 4.9.0 RC1 - LTO GCC 4.8.2 - LTO GCC 4.9.0 RC1 - Stock GCC 4.8.2 - Stock 4 8 12 16 20 SE +/- 0.06, N = 3 SE +/- 0.07, N = 3 SE +/- 0.01, N = 3 SE +/- 0.03, N = 3 13.75 13.80 13.80 13.99 -flto -flto 1. (CC) gcc options: -lavdevice -lavfilter -lavformat -lavcodec -lswresample -lswscale -lavutil -ldl -lasound -lSDL -lm -pthread -O3 -march=native -std=c99 -fomit-frame-pointer -fno-math-errno -fno-signed-zeros -fno-tree-vectorize -MMD -MF -MT
Open FMM Nero2D Total Time OpenBenchmarking.org Seconds, Fewer Is Better Open FMM Nero2D 2.0.2 Total Time GCC 4.8.2 - Stock GCC 4.8.2 - LTO 100 200 300 400 500 453.45 458.86 -flto 1. (CXX) g++ options: -O3 -march=native -lfftw3 -llapack -lf77blas -latlas -lm
Phoronix Test Suite v10.8.5