The current procedure for creating a new Datalad dataset is:
documented on https://portal.conp.ca/share:
datalad create <new_dataset_name>
cd <new_dataset_name>
datalad create-sibling-github <new_dataset_name> (this step will ask for a github username and password)
subsequent steps to initialise github repo documented on github:
4) git commit -m "first commit message"
5) git branch -M main (this step necessary since github changed the default master branch name to main)
6) git remote add origin https://github.com/<username>/<new_dataset_name>
documented on https://portal.conp.ca/share:
7) populate the dataset, for which there are multiple procedures
8) datalad save (can have commit message)
9) datalad push --to origin (or equivalent datalad push --to github)
In the case of the recrawled multicenter-phantom-test dataset, step 7 consists of:
a) load data into the new dataset directory
b) add README.txt and DATS.json (currently by copying from existing multicenter-phantom dataset)
c) edit .datalad/config to include the lines:
[credentials] force-ask = true
d) git add the files we do not want to archive (README, DATS.json, the edited datalad config file)
I have been using Shen's script at https://github.com/kongtiaowang/Loris-Crawler/ to carry out step 7a, and in the past couple of days have tried both:
having loris-crawler.py output directly to <new_dataset_name>
and:
having loris-crawler.py output to a local directory and then copying files to <new_dataset_name>
The end result, which so far as I can tell is indistinguishable for either of the above methods, is currently available in https://github.com/CONP-PCNO/multicenter-phantom-test.
Installing this repo using datalad install and then datalad get * proceeds very quickly, creates the dataset with links and correct structure, but does not appear to download any of the annexed files. My suspicion is that some of the manual setup procedure is duplicating or conflicting with commands run in loris-crawler.py.
The current procedure for creating a new Datalad dataset is:
documented on
https://portal.conp.ca/share:datalad create <new_dataset_name>cd <new_dataset_name>datalad create-sibling-github <new_dataset_name>(this step will ask for a github username and password)subsequent steps to initialise github repo documented on github:
4)
git commit -m "first commit message"5)
git branch -M main(this step necessary since github changed the defaultmasterbranch name tomain)6)
git remote add origin https://github.com/<username>/<new_dataset_name>documented on
https://portal.conp.ca/share:7) populate the dataset, for which there are multiple procedures
8)
datalad save(can have commit message)9)
datalad push --to origin(or equivalentdatalad push --to github)In the case of the recrawled
multicenter-phantom-testdataset, step 7 consists of:a) load data into the new dataset directory
b) add
README.txtandDATS.json(currently by copying from existing multicenter-phantom dataset)c) edit
.datalad/configto include the lines:[credentials] force-ask = trued)
git addthe files we do not want to archive (README, DATS.json, the edited datalad config file)I have been using Shen's script at https://github.com/kongtiaowang/Loris-Crawler/ to carry out step 7a, and in the past couple of days have tried both:
having
loris-crawler.pyoutput directly to<new_dataset_name>and:
having
loris-crawler.pyoutput to a local directory and then copying files to<new_dataset_name>The end result, which so far as I can tell is indistinguishable for either of the above methods, is currently available in https://github.com/CONP-PCNO/multicenter-phantom-test.
Installing this repo using
datalad installand thendatalad get *proceeds very quickly, creates the dataset with links and correct structure, but does not appear to download any of the annexed files. My suspicion is that some of the manual setup procedure is duplicating or conflicting with commands run inloris-crawler.py.