DistillSupra-0.2M


DistillSupra-0.2M is an ultra-compact causal language model with approximately 0.2 million parameters, produced by knowledge distillation from Supra-Mini-v4-2M.

It was trained 500 steps(1 Epoch) for 30 minutes on a GTX 750 Ti 4GB using generated text from the teacher.

The model was 10x compressed! That's crazy!

Architecture

Parameter Teacher Student
hidden_size 64 48
intermediate_size 128 96
num_hidden_layers 5 4
num_attention_heads 8 6
vocab_size 4096 4096
Parameters ~468k ~289k

Some outputs:

Prompt : Throughout history, great civilizations

Output: Throughout history, great civilizations to in, a be polrain for is with more the the be the for. of be of on (I.er The b M.A-R and or have that not is and the is this they, can for to to. is of a a, to ofs the for and the a. in the is to as of is that an that of and you the which is, the, for in be a are by’ of. and to a m

Prompt : The human brain is capable of

Output: The human brain is capable ofs in an more that in a new can is the this the a of the pS, the a to the other in not it... and with a to that be are of to for in of of ass. The be of the,.F-s be the of dLal. ins of be and of Sin: and or that a one that to and a a bFed, asRal., the, is a and as

Prompt : The most important principle in science is

The most important principle in science is a is a this are not for that the to of be digels-LC. to the in a the to, on to,

Why did supra created this trash?

We are currently researching knowledge distillation and this was the first step! Things will better up!

Benchmarks

Tasks Version Filter n-shot Metric Value Stderr
arc_easy 1 none 0 acc ↑ 0.2517 ± 0.0089
none 0 acc_norm ↑ 0.2677 ± 0.0091
blimp 2 none acc ↑ 0.5184 ± 0.0017
- blimp_adjunct_island 1 none 0 acc ↑ 0.7610 ± 0.0135
- blimp_anaphor_gender_agreement 1 none 0 acc ↑ 0.1920 ± 0.0125
- blimp_anaphor_number_agreement 1 none 0 acc ↑ 0.3650 ± 0.0152
- blimp_animate_subject_passive 1 none 0 acc ↑ 0.5790 ± 0.0156
- blimp_animate_subject_trans 1 none 0 acc ↑ 0.6450 ± 0.0151
- blimp_causative 1 none 0 acc ↑ 0.5210 ± 0.0158
- blimp_complex_NP_island 1 none 0 acc ↑ 0.5150 ± 0.0158
- blimp_coordinate_structure_constraint_complex_left_branch 1 none 0 acc ↑ 0.2970 ± 0.0145
- blimp_coordinate_structure_constraint_object_extraction 1 none 0 acc ↑ 0.3530 ± 0.0151
- blimp_determiner_noun_agreement_1 1 none 0 acc ↑ 0.5080 ± 0.0158
- blimp_determiner_noun_agreement_2 1 none 0 acc ↑ 0.4900 ± 0.0158
- blimp_determiner_noun_agreement_irregular_1 1 none 0 acc ↑ 0.4510 ± 0.0157
- blimp_determiner_noun_agreement_irregular_2 1 none 0 acc ↑ 0.5310 ± 0.0158
- blimp_determiner_noun_agreement_with_adj_2 1 none 0 acc ↑ 0.4730 ± 0.0158
- blimp_determiner_noun_agreement_with_adj_irregular_1 1 none 0 acc ↑ 0.4710 ± 0.0158
- blimp_determiner_noun_agreement_with_adj_irregular_2 1 none 0 acc ↑ 0.5370 ± 0.0158
- blimp_determiner_noun_agreement_with_adjective_1 1 none 0 acc ↑ 0.5220 ± 0.0158
- blimp_distractor_agreement_relational_noun 1 none 0 acc ↑ 0.4940 ± 0.0158
- blimp_distractor_agreement_relative_clause 1 none 0 acc ↑ 0.4800 ± 0.0158
- blimp_drop_argument 1 none 0 acc ↑ 0.6900 ± 0.0146
- blimp_ellipsis_n_bar_1 1 none 0 acc ↑ 0.2580 ± 0.0138
- blimp_ellipsis_n_bar_2 1 none 0 acc ↑ 0.2810 ± 0.0142
- blimp_existential_there_object_raising 1 none 0 acc ↑ 0.6180 ± 0.0154
- blimp_existential_there_quantifiers_1 1 none 0 acc ↑ 0.5550 ± 0.0157
- blimp_existential_there_quantifiers_2 1 none 0 acc ↑ 0.0750 ± 0.0083
- blimp_existential_there_subject_raising 1 none 0 acc ↑ 0.6750 ± 0.0148
- blimp_expletive_it_object_raising 1 none 0 acc ↑ 0.5820 ± 0.0156
- blimp_inchoative 1 none 0 acc ↑ 0.4010 ± 0.0155
- blimp_intransitive 1 none 0 acc ↑ 0.5550 ± 0.0157
- blimp_irregular_past_participle_adjectives 1 none 0 acc ↑ 0.2340 ± 0.0134
- blimp_irregular_past_participle_verbs 1 none 0 acc ↑ 0.7760 ± 0.0132
- blimp_irregular_plural_subject_verb_agreement_1 1 none 0 acc ↑ 0.4810 ± 0.0158
- blimp_irregular_plural_subject_verb_agreement_2 1 none 0 acc ↑ 0.4370 ± 0.0157
- blimp_left_branch_island_echo_question 1 none 0 acc ↑ 0.9310 ± 0.0080
- blimp_left_branch_island_simple_question 1 none 0 acc ↑ 0.3260 ± 0.0148
- blimp_matrix_question_npi_licensor_present 1 none 0 acc ↑ 0.0200 ± 0.0044
- blimp_npi_present_1 1 none 0 acc ↑ 0.6120 ± 0.0154
- blimp_npi_present_2 1 none 0 acc ↑ 0.5830 ± 0.0156
- blimp_only_npi_licensor_present 1 none 0 acc ↑ 0.0000 ± 0
- blimp_only_npi_scope 1 none 0 acc ↑ 0.0270 ± 0.0051
- blimp_passive_1 1 none 0 acc ↑ 0.5190 ± 0.0158
- blimp_passive_2 1 none 0 acc ↑ 0.5210 ± 0.0158
- blimp_principle_A_c_command 1 none 0 acc ↑ 0.6470 ± 0.0151
- blimp_principle_A_case_1 1 none 0 acc ↑ 1.0000 ± 0
- blimp_principle_A_case_2 1 none 0 acc ↑ 0.5290 ± 0.0158
- blimp_principle_A_domain_1 1 none 0 acc ↑ 1.0000 ± 0
- blimp_principle_A_domain_2 1 none 0 acc ↑ 0.5880 ± 0.0156
- blimp_principle_A_domain_3 1 none 0 acc ↑ 0.4790 ± 0.0158
- blimp_principle_A_reconstruction 1 none 0 acc ↑ 0.3470 ± 0.0151
- blimp_regular_plural_subject_verb_agreement_1 1 none 0 acc ↑ 0.5750 ± 0.0156
- blimp_regular_plural_subject_verb_agreement_2 1 none 0 acc ↑ 0.4710 ± 0.0158
- blimp_sentential_negation_npi_licensor_present 1 none 0 acc ↑ 1.0000 ± 0
- blimp_sentential_negation_npi_scope 1 none 0 acc ↑ 0.4020 ± 0.0155
- blimp_sentential_subject_island 1 none 0 acc ↑ 0.3030 ± 0.0145
- blimp_superlative_quantifiers_1 1 none 0 acc ↑ 0.5090 ± 0.0158
- blimp_superlative_quantifiers_2 1 none 0 acc ↑ 0.6150 ± 0.0154
- blimp_tough_vs_raising_1 1 none 0 acc ↑ 0.3080 ± 0.0146
- blimp_tough_vs_raising_2 1 none 0 acc ↑ 0.6980 ± 0.0145
- blimp_transitive 1 none 0 acc ↑ 0.5210 ± 0.0158
- blimp_wh_island 1 none 0 acc ↑ 0.4050 ± 0.0155
- blimp_wh_questions_object_gap 1 none 0 acc ↑ 0.9970 ± 0.0017
- blimp_wh_questions_subject_gap 1 none 0 acc ↑ 0.9990 ± 0.0010
- blimp_wh_questions_subject_gap_long_distance 1 none 0 acc ↑ 0.9980 ± 0.0014
- blimp_wh_vs_that_no_gap 1 none 0 acc ↑ 1.0000 ± 0
- blimp_wh_vs_that_no_gap_long_distance 1 none 0 acc ↑ 1.0000 ± 0
- blimp_wh_vs_that_with_gap 1 none 0 acc ↑ 0.0000 ± 0
- blimp_wh_vs_that_with_gap_long_distance 1 none 0 acc ↑ 0.0000 ± 0
wikitext 2 none 0 bits_per_byte ↓ 3.0125 ± N/A
none 0 byte_perplexity ↓ 8.0696 ± N/A
none 0 word_perplexity ↓ 70687.5638 ± N/A
Groups Version Filter n-shot Metric Value Stderr
blimp 2 none acc ↑ 0.5184 ± 0.0017

Final Thought

Knowledge distillation is a promising thing for us, we believe that LLMs can be helpful even being so small!

Downloads last month
197
Safetensors
Model size
289k params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SupraLabs/DistillSupra-0.2M

Finetuned
(1)
this model

Dataset used to train SupraLabs/DistillSupra-0.2M

Space using SupraLabs/DistillSupra-0.2M 1

Collection including SupraLabs/DistillSupra-0.2M