dartsort is a modular spike sorter built around a statistical clustering model and a new approach to probe motion. It is also a toolkit of modules for building spike sorters and other analyses of electrophysiology data.
We do not currently recommend DARTsort for production spike sorting purposes. Please feel free to open an issue or a discussion if you run into problems.
If you already have a Python environment with PyTorch working and you just want to install dartsort there, use
$ pip install dartsortIf you want to run the test suite or use dartsort.vis, you can install the optional dependencies with pip install dartsort[test,vis].
If you need to set up Python or PyTorch, I find that a conda-forge-based distribution is the most reliable at installing the GPU dependencies which PyTorch needs (note: conda-forge is different from the non-free Anaconda).
You can use conda-forge to install Python, dartsort, and its dependencies as follows:
- Follow the
conda-forgeinstallation instructions for your platform at https://conda-forge.org/download/ - Create an environment with
This will create an environment called
$ mamba env create -f environment.yml
dartsort, but you can change the name by adding-n othername. - Activate the environment:
$ mamba activate dartsort
- Install
dartsortand the rest of its dependencies by running the pip command above
Please read the important configuration details section below for information on the parameters of the DARTsortUserConfig.
In particular, you will most likely need to set the preprocessing flag.
For more detailed documentation of the main function parameters and configuration options, see the main API documentation.
dartsort can be run in Python with (for example):
import dartsort
# if the default settings make sense for you...
dartsort.dartsort(recording, output_dir)
# or, to set configuration options, something like...
dartsort_result = dartsort.dartsort(
recording,
output_dir,
cfg=dartsort.DARTsortUserConfig(
preprocessing="ibllike",
work_in_tmpdir=True,
copy_recording_to_tmpdir="yes",
),
)Here, recording is a SpikeInterface recording object (see their tutorial on reading various recording formats).
output_dir is the folder where dartsort will save its output.
There are more details on these and the rest of the arguments in the main API documentation.
Once you've run dartsort, you might want to check out the outputs and exporting section below.
Before running dartsort, please be aware of the following important configuration options.
preprocessing: dartsort doesn't preprocess your data by default (for now,preprocessing="none"by default), but this will change.- An easy choice is
preprocessing="ibllikecmr", which is a lightweight version of the IBL's strategy, replacing their highpass spatial filter with a common median reference. - If you've already preprocessed your data, use
"none"instead, or"standardize"if your preprocessing did not include a standardization step. - An implementation of the full IBL strategy is available as
"ibllike".
- An easy choice is
- The
copy_recording_to_tmpdircontrols whether the recording is copied to a scratch directory for faster reading. It can beTrue,False, or"if_preprocessing"(the default). By default the recording is cached ifpreprocessingis something other than"none".- (The
tmpdir_parentflag controls the scratch dir, which is your computer's default temporary folder if this is unset.)
- (The
do_motion_estimation=Trueby default, but you may like to disable it if you know that there is (say) less than 5 microns of total drift in your recording, or if you have handled this in your own preprocessing.work_in_tmpdircan be helpful in some cases where slow network drives are involved.
The dartsort_result = dartsort(...) function returns a dictionary dartsort_result containing a DARTsortSorting object: sorting = dartsort_result["sorting"].
This object has all the spike train data attached (as arrays under property names .times_samples and .times_seconds for spike times in samples and seconds, .labels for unit labels, and many others; print(sorting) to see some more).
If you already ran dartsort and want to load the output spike trains, use dartsort.load(output_dir) to get the DARTsortSorting object.
This object can also export itself to other formats:
- For a SpikeInterface
NumpySortingobject, usesorting.to_numpy_sorting() - For a Pynapple
TsGroup, usesorting.to_tsgroup() - To export to Phy, we currently suggest bridging through SpikeInterface. Start with
sorting.to_numpy_sorting()and follow the instructions in SpikeInterface's documentation for first creating aSortingAnalyzerand then exporting that to Phy. - For a simple dictionary of numpy arrays, use
dict = sorting.spike_feature_dict. - For pandas, use
sorting.to_pandas().
dartsort also saves motion information, returned as dartsort_result["motion"] or loaded after the fact as dartsort.try_load_motion_info(output_dir).
The data is saved to output_dir in the following files:
dartsort_sorting.npz: This NPZ file contains the final spike train under the keystimes_samples,channels, andlabels.matching1.h5: This HDF5 file contains spike features and other data from the last matching step. Amplitudes, localizations, and other features live in here; useh5lson the command line to see what's in there. Be aware that thelabelsdataset in this HDF5 is not the same as what's saved in thedartsort_sorting.npz.motion_info.pklis a pickledMotionInfoobject.- There may be a models/ folder containing PyTorch weights files with modeling quantities (for instance, featurization SVD bases, localization neural nets, Gaussian mixture model parameters). If you want to load these up, feel free to reach out for help.
To make some basic visualizations of the sorting result with matplotlib, try:
import dartsort.vis as dartvis
# gather outputs from dartsort
dartsort_result = dartsort(recording, output_dir, ...)
sorting = dartsort_result["sorting"]
motion = dartsort_result["motion"]
# or, if you already ran it
sorting = dartsort.load(output_dir)
motion = dartsort.try_load_motion_info(output_dir)
dartvis.visualize_sorting(
recording,
sorting,
vis_save_dir,
motion=motion,
make_unit_summaries=False,
)Set make_unit_summaries=True to create a summary plot for each unit.
Try running
$ dartsort -hon your command line to see usage instructions; parameters can be configured on the command line or read from a TOML file.
Please let us know if you run into any issues. If you feel that the issue is a software bug, feel free to open an issue or a discussion on GitHub. If it's more of a data-related or methodology thing, feel free to use the email on my GitHub profile.
- SpikeInterface
- DREDge on PyPI and the DREDge paper (PDF link).