Skip to content

README should include an example of a build command with explicit component paths #206

Description

@mikaelhg

System information

Describe the documentation issue

The README.md file should provide a build example for an arbitrary environment, which contains all of the required components, but not necessarily configured as the default choices. This is often the case with ML development workstations, which have multiple versions of CUDA, cuDNN, gcc, and friends, since different projects require the developer to use different versions for reproducibility reasons.

This example uses Bazelisk to manage multiple Bazel versions.

TF_CUDA_COMPUTE_CAPABILITIES will require #195 to be merged.

TF_CUDA_COMPUTE_CAPABILITIES=3.5,3.7,6.1,7.0,7.5 \
GCC_HOST_COMPILER_PATH=/usr/bin/gcc-8 \
CC=/usr/bin/gcc-8 \
CXX=/usr/bin/g++-8 \
TF_CUDA_PATHS=/usr/local/cuda-10.2.89,/usr/local/cudnn-10.2-7.6.5.32 \
CUDA_TOOLKIT_PATH=/usr/local/cuda-10.2.89 \
CUDNN_INSTALL_PATH=/usr/local/cudnn-10.2-7.6.5.32 \
USE_BAZEL_VERSION=3.1.0 \
TMP=/tmp \
mvn install -Dmaven.test.skip=true -Djavacpp.platform.extension=-gpu

We welcome contributions by users. Will you be able to update submit a PR (use the doc style guide) to fix the doc Issue?

No.

Activity

  1. rnett commented on Feb 3, 2021

    @rnett
    Contributor

    Since 99% of the native build is just building tensorflow, configuration should be handled in the same way (i.e. the configure script, or anything else). This is mentioned in #195, and clarified in the latest commit. Is there any reason that wouldn't work for you?

  2. mikaelhg commented on Feb 11, 2021

    @mikaelhg
    Author

    It's not that it's completely impossible to find out the practical information required to create a custom build, it's just that it would be incredibly easy to make finding this information convenient.

    As a result, hundreds or thousands of individual users wouldn't need to spend time searching for and learning this trivia, which they'll probably never need again, just to accomplish the task of building the binaries for compute capabilities other than 3.5 and 7.0.

  3. rnett commented on Feb 11, 2021

    @rnett
    Contributor

    You don't have to go hunting for any information though, you just clone tensorflow and run the configuration script. The reason we (and tensorflow) delegate to that is any hard-coded paths we provide will be wrong more often than they are right, so the script tries its best to autodetect them. Plus a lot of the configuration has been moved to the bazelrc files in the latest tensorflow release, so there should be even less required configuration.

    What exactly is so hard to find?

  4. karllessard commented on Feb 13, 2021

    @karllessard
    Collaborator

    It is possible though to simply use the TF archive that Bazel downloads and unzips just before building TF Java, instead of cloning the repo.

    I'm wondering if it would work to run the configure script from that unzipped archive during the Maven build directly, e.g. by passing a parameter like -Dnative.build.custom=true. I never tried running a blocking goal that requires user input in Maven though, is that possible?

    Or we simply add a new script in the Java repo that invoke Bazel to download the archive and run configure on it, bypassing our default .bazelrc config.

  5. rnett commented on Feb 13, 2021

    @rnett
    Contributor

    Or we simply add a new script in the Java repo that invoke Bazel to download the archive and run configure on it, bypassing our default .bazelrc config.

    I like this option much more. It would just be our own version of the configure script.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions