llm-lab
/

CLASP

Automatic Speech Recognition

Model card Files Files and versions Community

aboots commited on Apr 18

Commit

79624ff

verified ·

1 Parent(s): d622f31

Update README.md

Browse files

Files changed (1) hide show

README.md +21 -10

README.md CHANGED Viewed

@@ -7,10 +7,14 @@ datasets:
 pipeline_tag: automatic-speech-recognition
 ---
-[![arXiv](https://img.shields.io/badge/arXiv-Paper-<COLOR>.svg)](https://arxiv.org/abs/2412.13071) [![GitHub](https://img.shields.io/badge/GitHub-Code-181717?logo=github)](https://github.com/language-modeling-lab/CLASP)
 **CLASP** (Contrastive Language-Speech Pretraining) is a novel, lightweight, multilingual, multimodal representation designed for audio-text retrieval.
-To learn more about our proposed model, please refer to this [paper](https://arxiv.org/abs/2412.13071). All code is available on this [GitHub page](https://github.com/language-modeling-lab/CLASP).
 The newly introduced dataset, SpeechBrown, which we created for training this model, can be found on [this page](https://huggingface.co/datasets/llm-lab/SpeechBrown)
 CLASP creates powerful and meaningful semantic embeddings for raw speech in a 768-dimensional multilingual representation space. These embeddings can be used in various tasks such as speech retrieval or classification.
@@ -44,14 +48,21 @@ To use these models or train your own on custom datasets, please refer to our [G
 ## Citations
 If you find our paper, code, data, or models useful, please cite the paper:
 ```
-@misc{abootorabi2024claspcontrastivelanguagespeechpretraining,
-      title={CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval},
-      author={Mohammad Mahdi Abootorabi and Ehsaneddin Asgari},
-      year={2024},
-      eprint={2412.13071},
-      archivePrefix={arXiv},
-      primaryClass={cs.CL},
-      url={https://arxiv.org/abs/2412.13071},
 }
 ```

 pipeline_tag: automatic-speech-recognition
 ---
+[![arXiv](https://img.shields.io/badge/arXiv-Paper-<COLOR>.svg)](https://arxiv.org/abs/2412.13071) [![GitHub](https://img.shields.io/badge/GitHub-Code-181717?logo=github)](https://github.com/language-modeling-lab/CLASP) [![Website](https://img.shields.io/website?url=https%3A%2F%2Fmultimodalrag.github.io%2F)](https://clasp1.github.io/)
+[Models](https://huggingface.co/llm-lab/CLASP) | [Springer Link](https://link.springer.com/chapter/10.1007/978-3-031-88717-8_2) | [arXiv Link](https://arxiv.org/abs/2412.13071) | [Proposed Dataset](https://huggingface.co/datasets/llm-lab/SpeechBrown)  | [ACM Digital Library](https://dl.acm.org/doi/10.1007/978-3-031-88717-8_2) | [Website](https://clasp1.github.io/)
 **CLASP** (Contrastive Language-Speech Pretraining) is a novel, lightweight, multilingual, multimodal representation designed for audio-text retrieval.
+To learn more about our proposed model, please refer to this [paper](https://arxiv.org/abs/2412.13071), which is published at **ECIR 2025**. All code is available on this [GitHub page](https://github.com/language-modeling-lab/CLASP).
 The newly introduced dataset, SpeechBrown, which we created for training this model, can be found on [this page](https://huggingface.co/datasets/llm-lab/SpeechBrown)
 CLASP creates powerful and meaningful semantic embeddings for raw speech in a 768-dimensional multilingual representation space. These embeddings can be used in various tasks such as speech retrieval or classification.
 ## Citations
 If you find our paper, code, data, or models useful, please cite the paper:
 ```
+@inproceedings{10.1007/978-3-031-88717-8_2,
+                author = {Abootorabi, Mohammad Mahdi and Asgari, Ehsaneddin},
+                title = {CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval},
+                year = {2025},
+                isbn = {978-3-031-88716-1},
+                publisher = {Springer-Verlag},
+                address = {Berlin, Heidelberg},
+                url = {https://doi.org/10.1007/978-3-031-88717-8_2},
+                doi = {10.1007/978-3-031-88717-8_2},
+                abstract = {This study introduces CLASP (Contrastive Language-Speech Pretraining), a multilingual, multimodal representation tailored for audio-text information retrieval. CLASP leverages the synergy between spoken content and textual data. During training, we utilize our newly introduced speech-text dataset, which encompasses 15 diverse categories ranging from fiction to religion. CLASP’s audio component integrates audio spectrograms with a pre-trained self-supervised speech model, while its language encoding counterpart employs a sentence encoder pre-trained on over 100 languages. This unified lightweight model bridges the gap between various modalities and languages, enhancing its effectiveness in handling and retrieving multilingual and multimodal data. Our evaluations across multiple languages demonstrate that CLASP establishes new benchmarks in HITS@1, MRR, and meanR metrics, outperforming traditional ASR-based retrieval methods that rely on transcribing speech into text for subsequent text retrieval, especially in specific scenarios.},
+                booktitle = {Advances in Information Retrieval: 47th European Conference on Information Retrieval, ECIR 2025, Lucca, Italy, April 6–10, 2025, Proceedings, Part IV},
+                pages = {10–20},
+                numpages = {11},
+                keywords = {Multimodal IR, Speech Retrieval, Contrastive Learning},
+                location = {Lucca, Italy}
 }
 ```