Castor

This is the common repo for PyTorch deep learning models by the Data Systems Group at the University of Waterloo.

Models

Predictions Over One Input Text Sequence

For sentiment analysis, topic classification, etc.

Kim CNN: baseline convolutional neural network for sentence classification (Kim, EMNLP 2014)
conv-RNN: convolutional RNN (Wang et al., KDD 2017)

Predictions Over Two Input Text Sequences

For paraphrase detection, question answering, etc.

SM-CNN: Siamese CNN for ranking texts (Severyn and Moschitti, SIGIR 2015)
MP-CNN: Multi-Perspective CNN (He et al., EMNLP 2015)
NCE: Noise-Contrastive Estimation for answer selection applied on SM-CNN and MP-CNN (Rao et al., CIKM 2016)
VDPWI: Very-Deep Pairwise Word Interaction NNs for modeling textual similarity (He and Lin, NAACL 2016)
IDF Baseline: IDF overlap between question and candidate answers

Each model directory has a README.md with further details.

Setting up PyTorch

If you are an internal Castor contributor using GPU machines in the lab, follow the instructions here.

Castor is designed for Python 3.6 and PyTorch 0.4. PyTorch recommends Anaconda for managing your environment. We'd recommend creating a custom environment as follows:

$ conda create --name castor python=3.6
$ source activate castor

And installing the packages as follows:

$ conda install pytorch torchvision -c pytorch

Other Python packages we use can be installed via pip:

$ pip install -r requirements.txt

Code depends on data from NLTK (e.g., stopwords) so you'll have to download them. Run the Python interpreter and type the commands:

>>> import nltk
>>> nltk.download()

Finally, run the following inside the utils directory to build the trec_eval tool for evaluating certain datasets.

$ ./get_trec_eval.sh

Data and Pre-Trained Models

If you are an internal Castor contributor using GPU machines in the lab, follow the instructions here.

To fully take advantage of code here, clone these other two repos:

Castor-data: embeddings, datasets, etc.
Caster-models: pre-trained models

Organize your directory structure as follows:

.
├── Castor
├── Castor-data
└── Castor-models

For example (using HTTPS):

$ git clone https://github.com/castorini/Castor.git
$ git clone https://git.uwaterloo.ca/jimmylin/Castor-data.git
$ git clone https://git.uwaterloo.ca/jimmylin/Castor-models.git

After cloning the Castor-data repo, you need to unzip embeddings and run data pre-processing scripts. You can choose to follow instructions under each dataset and embedding directory separately, or just run the following script in Castor-data to do all of the steps for you:

$ ./setup.sh

Name		Name	Last commit message	Last commit date
Latest commit History 109 Commits
anserini_dependency		anserini_dependency
common		common
conv_rnn		conv_rnn
datasets		datasets
docs		docs
idf_baseline		idf_baseline
kim_cnn		kim_cnn
mp_cnn		mp_cnn
nce		nce
sm_cnn		sm_cnn
utils		utils
vdpwi		vdpwi
.gitignore		.gitignore
MyFirstColabNotebook.ipynb		MyFirstColabNotebook.ipynb
README.md		README.md
__init__.py		__init__.py
baseline_results.tsv		baseline_results.tsv
config.cfg.example		config.cfg.example
requirements.txt		requirements.txt
run_ui.sh		run_ui.sh
setup.py		setup.py

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

Castor

Models

Predictions Over One Input Text Sequence

Predictions Over Two Input Text Sequences

Setting up PyTorch

Data and Pre-Trained Models

About

Releases

Packages

Languages

marinecarpuat/Castor

Folders and files

Latest commit

History

Repository files navigation

Castor

Models

Predictions Over One Input Text Sequence

Predictions Over Two Input Text Sequences

Setting up PyTorch

Data and Pre-Trained Models

About

Resources

Stars

Watchers

Forks

Releases

Packages 0

Languages

Packages