Update README.md
Browse files
README.md
CHANGED
@@ -16,7 +16,7 @@ One of the unique features of our approach to adaptation lies in the fact that,
|
|
16 |
|
17 |
An intriguing aspect of adapting T-pro-it-1.0 is that this model was obtained through continuous pretraining on over 100 billion tokens of Russian-language data using full fine-tuning. Despite this extensive prior training, our methodology still worked effectively (note: the original base model Qwen2.5-32B was adapted!), and the resulting adapted version either outperformed or matched T-pro-it-1.0 on several benchmarks. Moreover, it demonstrated higher efficiency in Russian-language tokenization.
|
18 |
|
19 |
-

|
|
|
16 |
|
17 |
An intriguing aspect of adapting T-pro-it-1.0 is that this model was obtained through continuous pretraining on over 100 billion tokens of Russian-language data using full fine-tuning. Despite this extensive prior training, our methodology still worked effectively (note: the original base model Qwen2.5-32B was adapted!), and the resulting adapted version either outperformed or matched T-pro-it-1.0 on several benchmarks. Moreover, it demonstrated higher efficiency in Russian-language tokenization.
|
18 |
|
19 |
+

|
20 |
|
21 |
## Papers
|
22 |
Tikhomirov M., Chernyshov D. Facilitating Large Language Model Russian Adaptation with Learned Embedding Propagation //Journal of Language and Education. β 2024. β Π’. 10. β β. 4. β Π‘. 130-145. (Preprint: https://arxiv.org/abs/2412.21140)
|