Skip to content

Feature/tame run times - #1

Merged
oleksandr-pavlyk merged 4 commits into
masterfrom
feature/tame-run-times
Sep 9, 2017
Merged

oleksandr-pavlyk merged 4 commits into
masterfrom
feature/tame-run-times

Conversation

@oleksandr-pavlyk

Copy link
Copy Markdown
Contributor
  1. Adjusted architecture flag when compiling native C benchmarks for FFT to use COMMON-AVX512, rather than CORE-AVX512. See reference.

  2. Default choice of 16 runs per batch with 24 repetitions of batch runs, combined with slow performance of FFT in stock NumPy/SciPy, results is excessive run-times for benchmarks of open source stack. Therefore, lowered default 24 repetitions to 6, and added code that would dynamically lower batch size if the batch run time exceeds 5, and rescale the timing to what it would have been should the request batch size was used.

@anton-malakhov @rscohn2

CORE-AVX512 -> COMMON-AVX512 for consistency with Windows and for better performance on KNL
During the warm-up, measure time to execute one call, check if total number of running the batch exceeds 5 seconds. If so, reduce the actual batch_size used and rescale the result, extrapolating the time for the requested batch size.

This combats excessive benchmark run-times when using stock Python.
This was caused by setting DFTI_COMPLEX_STORAGE to DFTI_COMPLEX_COMPLEX
@oleksandr-pavlyk
oleksandr-pavlyk merged commit a95ce7b into master Sep 9, 2017
@oleksandr-pavlyk
oleksandr-pavlyk deleted the feature/tame-run-times branch September 9, 2017 14:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant