Benchmark data in "Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions".
AI & ML interests
LLM reasoning
Recent Activity
View all activity
models 7
Miaow-Lab/Qwen3.5-9B-decision-subtrajectory-32k-epoch3
Text Generation • 9B • Updated • 255
Miaow-Lab/Qwen3.5-9B-decision-subtrajectory-32k-epoch2
Text Generation • 9B • Updated • 274
Miaow-Lab/Qwen3.5-9B-decision-subtrajectory-32k-epoch1
Text Generation • 9B • Updated • 279
Miaow-Lab/RLVR-Linearity-Checkpoints
Text Generation • Updated
Miaow-Lab/STT-Agent-RL
196k • Updated • 18 • 1
Miaow-Lab/STT-Agent-SFT
196k • Updated • 7 • 1
Miaow-Lab/SSAE-Checkpoints
Feature Extraction • Updated
datasets 6
Miaow-Lab/v7-kernel-candidates
Viewer • Updated • 5.16k
Miaow-Lab/STT-Arena
Preview • Updated • 74 • 2
Miaow-Lab/OpenSkillRisk
Updated • 84
Miaow-Lab/RUT-Bench
Viewer • Updated • 1.64k • 31 • 1
Miaow-Lab/RLVR-Linearity-Dataset
Viewer • Updated • 40.3k • 46
Miaow-Lab/SSAE-Dataset
Viewer • Updated • 1.28M • 25