Skip to content

Rewrite native FFT benchmarks - #6

Merged
oleksandr-pavlyk merged 22 commits into
IntelPython:masterfrom
bibikar:simplify_native
Jan 13, 2020
Merged

oleksandr-pavlyk merged 22 commits into
IntelPython:masterfrom
bibikar:simplify_native

Conversation

@bibikar

@bibikar bibikar commented Jan 10, 2020 •

Copy link
Copy Markdown
Contributor

@oleksandr-pavlyk

This PR includes a complete rewrite of native FFT benchmarks using MKL DFTI:

  • Unify all benchmarks into one file
  • Add support for single-precision dtypes and 4-, 5-, 6-, 7-dimensional FFTs
  • Switch to command-line options (getopt_long) instead of env
  • Add help message

Copying superfluous harmonics for conjugate-even outputs produced by the FFT for real inputs as is done in mkl_fft is not yet implemented, but using the --rfft argument allows skipping that step for now. I don't think our native benchmarks have ever supported FFTs of real inputs with full outputs.

The benchmark still works on Windows, but only short options are supported there.

usage: ./fft_bench [args] size
Benchmark FFT using Intel(R) MKL DFTI.

FFT problem arguments:
  -t, --threads=THREADS    use THREADS threads for FFT execution
                           (default: use MKL's default)
  -d, --dtype=DTYPE        use DTYPE as the FFT domain. For a list of
                           understood dtypes, use '-d help'.
                           (default: complex128)
  -r, --rfft               do not copy superfluous harmonics when FFT
                           output is even-conjugate, i.e. for real inputs
  -P, --in-place           allow overwriting the input buffer with the
                           FFT outputs
  -c, --cached             use the same DFTI descriptor for the same
                           outer loop, i.e. "cache" the descriptor

Timing arguments:
  -i, --inner-loops=IL     time the benchmark IL times for each printed
                           measurement. Copies are not included in the
                           measurements. (default: 16)
  -o, --outer-loops=OL     print OL measurements. (default: 5)

Output arguments:
  -p, --prefix=PREFIX      output PREFIX as the first value in outputs
                           (default: 'Native-C')
  -H, --no-header          do not output CSV header. This can be useful
                           if running multiple benchmarks back-to-back.
  -h, --help               print this message and exit

The size argument specifies the input matrix size as a tuple of positive
decimal integers, delimited by any non-digit. For example, both
(101, 203, 305) and 101x203x305 denote the same 3D FFT.

Comment thread native/Makefile Outdated
CFLAGS = -m64 -fPIC -fp-model strict -O3 -g -fomit-frame-pointer \
-DNDEBUG -qopenmp -xSSE4.2 -axCORE-AVX2,COMMON-AVX512 \
-lmkl_rt
-lmkl_rt -DDEBUG -Wall -pedantic

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What is the point of having -DDEBUG alongside -DNDEBUG ?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I had written some debug messages in an earlier revision enabled with -DDEBUG, but they've been since removed, so this flag is now useless. I went ahead and removed it - thanks!

@oleksandr-pavlyk

Copy link
Copy Markdown
Contributor

Help seems to suggest that option -i and -o are valid, yet:

(t_scipy-1.4.0) [10:14:08 vmlin native]$ ./fft_bench -t 1 -d float32 -P -c --inner-loops=10 --outer-loops=7 7  # works as advertised
prefix,function,threads,dtype,size,place,cached,time
Native-C,fft,1,float32,7,in-place,cached,0.0092834
Native-C,fft,1,float32,7,in-place,cached,1.9109e-05
Native-C,fft,1,float32,7,in-place,cached,1.8982e-05
Native-C,fft,1,float32,7,in-place,cached,1.9297e-05
Native-C,fft,1,float32,7,in-place,cached,1.8928e-05
Native-C,fft,1,float32,7,in-place,cached,1.8946e-05
Native-C,fft,1,float32,7,in-place,cached,1.8844e-05
(t_scipy-1.4.0) [10:14:15 vmlin native]$ ./fft_bench -t 1 -d float32 -P -c --i=10 --outer-loops=7 7
./fft_bench: option '--i=10' is ambiguous; possibilities: '--inner-loops' '--in-place'
(t_scipy-1.4.0) [10:14:21 vmlin native]$ ./fft_bench -t 1 -d float32 -P -c -i=10 --outer-loops=7 7
./fft_bench: invalid option -- 'i'  # I was expecting this to actually work

@bibikar

bibikar commented Jan 10, 2020

Copy link
Copy Markdown
Contributor Author

Short options -ioH were missing from the getopt_long call. I added them, but instead of -i=10, it is required to specify -i10 or -i 10.

@oleksandr-pavlyk oleksandr-pavlyk left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks @bibikar

@oleksandr-pavlyk

Copy link
Copy Markdown
Contributor

👍

@oleksandr-pavlyk
oleksandr-pavlyk merged commit c8e0913 into IntelPython:master Jan 13, 2020
@bibikar

bibikar commented Jan 13, 2020

Copy link
Copy Markdown
Contributor Author

Thanks!

@bibikar
bibikar deleted the simplify_native branch January 23, 2020 05:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants