Skip to content

Commit 6e49cb3

Browse files
committed
numpy: download the documentation files automatically
1 parent bafe456 commit 6e49cb3

2 files changed

Lines changed: 33 additions & 9 deletions

File tree

‎docs/file-scrapers.md‎

Lines changed: 2 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -149,11 +149,8 @@ mv man7.org/linux/man-pages/ docs/man/
149149

150150
## NumPy
151151

152-
```sh
153-
mkdir --parent docs/numpy~$VERSION/; \
154-
curl https://numpy.org/doc/$VERSION/numpy-html.zip | \
155-
bsdtar --extract --file=- --directory=docs/numpy~$VERSION/
156-
```
152+
Nothing to do — the scraper downloads and extracts the HTML archive into
153+
`docs/numpy~$VERSION` automatically when it's missing.
157154

158155
## OpenGL
159156

‎lib/docs/scrapers/numpy.rb‎

Lines changed: 31 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -1,8 +1,4 @@
11
module Docs
2-
# Requires downloading the documents to local disk first.
3-
# Go to https://numpy.org/doc/, click "HTML+zip" to download
4-
# (example url: https://numpy.org/doc/2.4/numpy-html.zip),
5-
# then extract into "docs/numpy~#{version}/"
62
class Numpy < FileScraper
73
self.name = 'NumPy'
84
self.type = 'sphinx'
@@ -137,5 +133,36 @@ class Numpy < FileScraper
137133
def get_latest_version(opts)
138134
get_latest_github_release('numpy', 'numpy', opts)
139135
end
136+
137+
private
138+
139+
# The documents predating 1.18 were published on docs.scipy.org instead.
140+
def archive_url
141+
@archive_url ||= [
142+
"https://numpy.org/doc/#{self.class.version}/numpy-html.zip",
143+
"https://docs.scipy.org/doc/numpy-#{self.class.release}/numpy-html-#{self.class.release}.zip"
144+
].find { |url| Request.run(url, method: :head).success? }
145+
end
146+
147+
def download_source
148+
require 'unix_utils'
149+
150+
raise SetupError, "No documentation archive found for NumPy #{self.class.release}." if archive_url.nil?
151+
152+
instrument 'info.doc', msg: %(Downloading #{archive_url}...)
153+
archive = UnixUtils.curl(archive_url)
154+
155+
instrument 'info.doc', msg: %(Extracting the documentation files to "#{source_directory}"...)
156+
directory = UnixUtils.unzip(archive)
157+
158+
FileUtils.mkpath(File.dirname(source_directory))
159+
# The archive holds the whole documentation, of which the older versions
160+
# are scraped from the "reference" subdirectory alone.
161+
root = base_url.path.end_with?('/reference/') ? File.join(directory, 'reference') : directory
162+
FileUtils.mv(root, source_directory)
163+
ensure
164+
FileUtils.rm_f(archive) if archive
165+
FileUtils.rm_rf(directory) if directory
166+
end
140167
end
141168
end

0 commit comments

Comments
 (0)