Skip to content

Latest commit

 

History

History
9 lines (6 loc) · 408 Bytes

File metadata and controls

9 lines (6 loc) · 408 Bytes

doc-classification

#TO-DO

  1. Currently the scraper for guardian.com extracts articles for the month of MAY, make it a command line argument for the user to specify the date, month and year.

  2. Make a new version to scrape all the articles from the begining.

  3. Combine create and build dataset into one script

Industry wise/ company wise extraction of pdfs from annualreports.com on a seperate repo.