Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🛠️
Working on SLM Arena
37.2
TFLOPS
CodeSoft
PRO
CodeSoft
9
3
27
Follow
wayneworkman2012's profile picture
arthurblg1802's profile picture
MultivexAI's profile picture
25 followers
·
119 following
https://codesft.dev
codesftdev
AI & ML interests
Working on small language models
Recent Activity
liked
a model
about 5 hours ago
DALabCommunity/Haidass1.5-143M
upvoted
a
paper
about 5 hours ago
On-Policy Self-Distillation in Diffusion Models
replied
to
their
post
about 5 hours ago
Over the past week or so, I've been working on some models, those released being https://huggingface.co/CodeSoft/sorbet-25m and https://huggingface.co/CodeSoft/sorbet-v2-25m. In general, I'm a little confused because no matter what hyperparameters I change or datasets I add/remove, the benchmarks never move up. In a recent project, where I attached a TN-gram block to Sorbet-v2-25M, it still stayed the same on benchmarks despite the TN-gram clearly learning (due to the perplexity being lower with the TN-gram attached). When I changed the corpus to favor higher density text (the first paragraphs of Wikipedia articles and synthetic math), the benchmarks either stayed flat or went down. Does anyone have ideas on what I can do to improve my models? I'd really appreciate any feedback!
View all activity
Organizations
CodeSoft
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
liked
a model
about 5 hours ago
DALabCommunity/Haidass1.5-143M
0.1B
•
Updated
about 5 hours ago
•
7
liked
a Space
1 day ago
Running
133
Microduck Sandbox
🐤
133
Showcase a bold, stylized landing page with fancy titles
liked
a model
3 days ago
Qwen/Qwen3.8-Flash-Next
Image-Text-to-Text
•
180B
•
Updated
2 days ago
•
4.81k
•
4.18k
liked
a model
4 days ago
SupraLabs/Supra2-Medium-Instruct
Text Generation
•
25.4M
•
Updated
8 days ago
•
1.2k
•
14
liked
a model
6 days ago
AxionML/Qwen3.5-0.8B-Base-NVFP4
Image-Text-to-Text
•
0.7B
•
Updated
Mar 3
•
149
•
1
liked
a model
7 days ago
AuroraAI-Research/Aurora-80K
Text Generation
•
Updated
9 days ago
•
426
•
22
liked
2 models
10 days ago
allenai/OLMo-2-0425-1B
Text Generation
•
1B
•
Updated
May 28, 2025
•
695k
•
81
bartowski/Ling-3.0-tiny-GGUF
Text Generation
•
8B
•
Updated
11 days ago
•
20.1k
•
24
liked
2 models
11 days ago
fromziro/MrPong
Reinforcement Learning
•
28.5k
•
Updated
3 days ago
•
216
•
12
empero-ai/Qwen3.8-27B-Ridge-GGUF
Image-Text-to-Text
•
27B
•
Updated
13 days ago
•
212k
•
286
liked
a model
13 days ago
empero-ai/Qwen3.8-9B-Distill-GGUF
Text Generation
•
9B
•
Updated
13 days ago
•
230k
•
184
liked
a model
14 days ago
inclusionAI/Ling-3.0-tiny-fp8
8B
•
Updated
10 days ago
•
6.71k
•
29
liked
a model
15 days ago
Qwen/Qwen3.8-27B
Image-Text-to-Text
•
28B
•
Updated
15 days ago
•
3.46M
•
•
13.2k
liked
a model
17 days ago
inclusionAI/Ling-3.0-tiny
Text Generation
•
8B
•
Updated
10 days ago
•
17.3k
•
374
liked
a model
19 days ago
BananaMind/BananaMind-2-Pro-Preview
Text Generation
•
0.2B
•
Updated
12 days ago
•
1.04k
•
24
liked
a Space
26 days ago
Running
88
Open SLM Leaderboard
🏆
88
Open Small Language Model Leaderboard
liked
a model
28 days ago
0xSero/DeepSeek-V4-Flash-0731-REAP
Text Generation
•
193B
•
Updated
29 days ago
•
5.86k
•
34
liked
a dataset
about 1 month ago
empero-ai/gpt-5.6-luna-sft-900x
Viewer
•
Updated
Jul 16
•
890
•
233
•
6
liked
a model
about 1 month ago
moonshotai/Kimi-K3
Image-Text-to-Text
•
2.8T
•
Updated
9 days ago
•
2.68M
•
•
11.1k
liked
a dataset
about 1 month ago
SupraLabs/LLM-self-identification
Viewer
•
Updated
about 1 month ago
•
459
•
273
•
12
Load more