LightDec_Arthur

LightDec_Arthur is a small, fast, calibrated decision model from Falcons.ai. Given a piece of state (an email, a ticket, an agent trace, a log, a JSON record) and a typed question, it picks an answer from a closed set of options in a single forward pass and returns calibrated probabilities, so your application knows when to trust the answer and when to defer to a human or a larger model.

Unlike Falconsai/LightDec_V2, it has no pretrained backbone and no tokenizer. It reads raw UTF-8 bytes and was trained from scratch on the FalconDec_V2 decision mixture. At about 12 million parameters it is roughly 13× smaller than LightDec_V2, trading accuracy for size and speed.

Architecture Arthur (byte-level recursive encoder + option-interaction head), trained from scratch
Backbone None
Parameters 12,268,033 (fp16 weights; model folder 24.5 MB)
Context 1,024 tokens of 4 bytes (4,096 bytes) per call; 2,048 tokens for questions with more than 24 options
Question types choice, noul (yes/no), score (ordinal levels)
Options per question 2 or more in one pass (up to 77 in the test set)
Test accuracy 59.6% micro / 58.4% macro over 26,597 test items, 58 tasks
Calibration ECE 0.030 after temperature scaling (0.246 before)
Latency 5.1 ms per decision on GPU (NVIDIA RTX PRO 6000 Blackwell); CPU not yet measured
Version Arthur notebook v0.8.0, base tier

What it does

LightDec_Arthur answers the same kind of small, structured questions as LightDec: Which queue does this ticket go to? Is this code a clone of that one? Is this prompt a jailbreak? Which intent is this request? It is not a generative model and never produces free text; it only chooses among the options you give it.

Three question types are supported:

  • choice: pick one of N options.
  • noul: a yes/no judgment (options Yes, No).
  • score: an ordinal scale such as ["Calm", "Annoyed", "Furious"], trained with a ranked-probability term so near misses cost less than far ones.

The calibrated top probability works as a confidence. With a threshold of 0.80, the model answered 21.5% of test questions on its own and was 89.7% accurate on those; 94.6% of its mistakes fell in the deferred remainder.

How it works

The question, the options and the state are packed into one byte sequence:

[CLS] question [SEP] [OPT] option_1 [OPT] option_2 ... [OPT] option_k [SEP] state [SEP]

Each byte gets an embedding, plus hashed embeddings of the 2-, 3- and 4-byte fragments ending at it (65,536 buckets). A depthwise convolution mixes neighbouring bytes, and every 4 bytes are pooled into one position; each [OPT] marker starts a new position. One transformer block (8 heads, rotary positions) is applied 6 times over the text, with a learned step vector added before each pass.

The vector at each option's [OPT] position, plus the mean of the text and a question-type embedding, forms the option vector. The same block is applied twice more over the option vectors, without position encoding, so options can compare themselves with each other. A small MLP gives one logit per option. A temperature per question type and option-count bucket (2, 3–5, 6–12, 13+ options), fitted on validation data and stored in config.json, calibrates the softmax.

Usage

This repository does not yet include a standalone loader. The model definition (Arthur), the byte layout (assemble) and the temperature step live in the Arthur training notebook. To score a decision with the notebook's code:

import json, numpy as np, torch
from safetensors.torch import load_file

cfg = json.load(open("config.json"))                      # tier config, byte layout, temperatures
net = Arthur(TIER_CFG["base"])                             # class from the Arthur notebook
net.load_state_dict({k: v.float() for k, v in load_file("model.safetensors").items()})
net.to(DEVICE).eval()

d = dict(state="I was billed twice for my Pro plan in March. Please refund the duplicate charge.",
         question="What is this support email primarily about?",
         options=["billing", "technical", "feature request", "account"], qtype="choice")
ids, win, wmask, qt = batch_tensors([d])                   # from the Arthur notebook
with torch.no_grad():
    z = net(ids, win, wmask, qt)[0, :len(d["options"])].float().cpu().numpy()
t = np.array(cfg["temperatures"])[QTYPES[d["qtype"]], bucket(len(d["options"]))]
p = np.exp((z - z.max()) / t); p /= p.sum()
print(d["options"][int(p.argmax())], round(float(p.max()), 3), "DEFER" if p.max() < 0.8 else "")

Evaluation

All numbers come from results_full.json, written by the training run. The test split has 26,597 items across 58 tasks in 9 domains. Tasks marked held-out have no training or validation data.

By domain

Accuracy is the mean over the domain's tasks.

Domain Tasks Test items Accuracy
support 2 1,000 99.5%
code 9 3,854 74.1%
guardrails 4 1,294 69.0%
agentic 4 925 60.5%
workflows 4 2,000 60.4%
intents 4 1,800 58.3%
policy 9 4,500 56.5%
reasoning 18 9,224 47.2%
classification 4 2,000 43.0%

Head-to-head against Arthur nano

Arthur nano is the 1.6 MB tier of the same architecture (793,729 parameters), scored on the same test set. It was trained for 8 epochs with a 512-token window, so this comparison mixes model size, epochs and input length.

On the 58 shared tasks, LightDec_Arthur averaged 58.4% against 53.6% for Arthur nano, winning on 39 tasks, tying on 3 and losing on 16.

Per-task head-to-head
Task n Arthur Arthur nano Δ
policy/return_window_transfer 500 82.6% 31.4% +51.2
mbpp/solution 500 79.2% 49.2% +30.0
codexglue/doc_to_code 504 88.5% 63.7% +24.8
qasc/mcq 500 46.8% 22.4% +24.4
codexglue/code_to_doc 504 83.9% 64.9% +19.1
codexglue/func_name 467 69.8% 51.8% +18.0
hotpotqa/retrieve 497 58.6% 41.6% +16.9
hotpotqa/comparison_yes_no 26 65.4% 50.0% +15.4
sciq/mcq 498 60.8% 47.6% +13.3
banking77/intent 500 58.2% 46.4% +11.8
massive_en/intent 500 85.4% 75.6% +9.8
policy/refund_approval_transfer 500 52.0% 42.6% +9.4
snli/nli 988 72.4% 63.8% +8.6
snli/contradicts 500 85.4% 77.4% +8.0
policy/invoice_overdue_transfer 500 48.8% 41.0% +7.8
snli/must_be_true 500 90.8% 83.6% +7.2
policy/sla_urgency_transfer 500 30.2% 23.0% +7.2
openbookqa/mcq 500 36.2% 30.2% +6.0
policy/table_extreme_transfer 500 28.4% 22.4% +6.0
arc_challenge/mcq 500 29.4% 23.8% +5.6
winogrande/blank 500 56.2% 50.8% +5.4
sst5/score 500 27.4% 22.6% +4.8
clinc150/intent 500 84.8% 80.2% +4.6
policy/free_shipping_transfer 500 80.0% 75.6% +4.4
humaneval/completion 119 42.0% 37.8% +4.2
mnli/claim 500 52.0% 48.6% +3.4
typed_decisions/agent_trace_observability 500 69.8% 66.6% +3.2
commonsense_qa/mcq 493 19.1% 16.8% +2.2
boolq/yes_no 500 62.6% 60.4% +2.2
arc_easy/mcq 500 30.4% 28.4% +2.0
jailbreak/detect 262 91.2% 89.3% +1.9
devign/vulnerability 500 64.2% 62.4% +1.8
hellaswag/continuation 500 28.0% 26.2% +1.8
banking77/intent_77 300 4.7% 3.0% +1.7
civil_comments/toxic 500 86.8% 85.6% +1.2
ag_news/topic 500 84.6% 83.6% +1.0
bitext/route 500 99.4% 98.8% +0.6
bigclonebench/clone 500 95.4% 95.0% +0.4
mbpp/bugspot 256 43.8% 43.4% +0.4
codexglue/lang_id 504 99.8% 99.8% +0.0
policy/access_control_transfer 500 100.0% 100.0% +0.0
policy/invoice_total_transfer 500 52.0% 52.0% +0.0
scitail/support 500 66.0% 66.2% -0.2
bitext/category 500 99.6% 99.8% -0.2
mmlu/mcq 500 28.4% 28.8% -0.4
gsm8k/math 500 30.6% 31.2% -0.6
yelp/score 500 48.6% 50.6% -2.0
typed_decisions/security_incidents 500 65.8% 68.2% -2.4
typed_decisions/invoice_processing 500 56.0% 59.4% -3.4
policy/count_threshold_transfer 500 34.8% 38.2% -3.4
agentharm/refuse 416 50.7% 54.3% -3.6
aqua_rat/math 247 21.9% 27.5% -5.7
counsel/step_has_error 201 72.1% 78.1% -6.0
anli/nli 498 32.1% 38.1% -6.0
counsel/critique_quality 201 45.8% 53.2% -7.5
emotion/6way 500 11.6% 19.2% -7.6
typed_decisions/customer_service 500 50.0% 58.0% -8.0
prompt_injections/detect 116 47.4% 56.0% -8.6

Input length

At the 1,024-token window, 385 of 972,625 training and test decisions (0.04%) have text cut off, concentrated in counsel/critique_quality (7.3%), jailbreak/detect (7.0%) and counsel/step_has_error (4.6%).

Decision Index 0.2 (partial)

LightDec_Arthur was scored on 39 benchmarks of the public Decision Index 0.2 with the index kit's own runner and scorers. Five benchmarks were not built (ToolRet and BRIGHT need tiktoken, POP909 needs mido, ChessBench needs python-chess, HLE is gated), so this is not a Decision Index score. The kit's rules apply: nothing is cut, and a request longer than 8,192 bytes, or a question of a type the harness does not map (score), is Unsupported and counts as wrong.

On 35 counted benchmarks, 21 score exactly 0 skill (chance level), and the best is SGD/SGD-X at 8.84. Five benchmarks share a source with the training mixture (ANLI, HellaSwag, WinoGrande, GSM8K, CLINC150), yet score near chance: the model learned those datasets in the training mixture's question format and does not transfer to the index's formats. For example, CLINC150 scores 84.8% on this card's test but 0.19 skill on the index, which asks all 151 intents plus out-of-scope in one question. 140 training decisions share text with index test rows (74 HellaSwag, 26 MMLU, 25 MMLU-Pro, 10 ARC-Easy, 4 ARC-Challenge, 1 BBH).

Per-benchmark results (39 benchmarks)
# Benchmark Skill (0 = chance, 100 = perfect) Raw Coverage Unsupported Counted in index
#1 BFCL 2.94 0.2810 0.9799 34 yes
#3 API-Bank 0.00 0.0000 0.0000 508 yes
#4 BANKING77 0.00 0.0014 1.0000 0 yes
#5 CLINC150+OOS 0.19 0.0079 1.0000 0 yes
#6 RouterBench 0.00 0.5304 0.9266 367 yes
#9 Home appliance simulator 0.00 0.0000 1.0000 0 yes
#10 SGD/SGD-X 8.84 0.4521 0.9996 1 yes
#11 ContractNLI 0.00 0.0833 0.3659 78 yes
#12 ANLI 0.30 0.3344 1.0000 0 yes
#20 BPoMP 1.61 0.5081 1.0000 0 yes
#21 Humicroedit 0.00 0.4920 1.0000 0 yes
#23 cfcolor 0.00 0.4725 1.0000 0 yes
#24 MMLU 0.73 0.2555 1.0000 0 shown only
#25 GPQA Diamond 1.36 0.2602 1.0000 0 yes
#26 ARC-Easy 2.39 0.2681 1.0000 0 shown only
#27 ARC-Challenge 3.74 0.2782 1.0000 0 shown only
#28 WinoGrande 0.72 0.5036 1.0000 0 yes
#29 HellaSwag 3.25 0.2744 1.0000 0 yes
#30 GSM8K 0.52 0.1778 1.0000 0 yes
#32 MuSR 0.00 0.3577 1.0000 0 yes
#33 SATA-Bench 0.00 0.0012 0.9594 67 yes
#34 SimpleBench 0.00 0.0000 1.0000 0 shown only
#37 Amazon ESCI 0.00 0.1965 1.0000 0 yes
#38 ACOS 0.00 0.0000 1.0000 0 yes
#39 FinEntity 0.00 0.2199 1.0000 0 yes
#40 iSarcasmEval 5.16 0.2628 1.0000 0 yes
#41 VAST 0.00 0.3276 1.0000 0 yes
#42 NLI4CT 0.22 0.4871 0.9784 119 yes
#43 CRUXEval 7.87 0.4193 1.0000 0 yes
#44 CLadder 0.00 0.4976 1.0000 0 yes
#48 ForecastBench 0.00 0.0000 0.9853 149 yes
#50 Habermas Machine 0.00 0.2894 1.0000 0 yes
#56 PhishNChips phishing decisions 0.00 0.4930 1.0000 0 yes
#57 MMLU-Pro 0.02 0.1111 1.0000 0 yes
#58 BBH fixed-option tasks 0.14 0.3111 1.0000 0 yes
#59 RAGTruth response-level hallucination 0.00 0.2548 0.9837 44 yes
#61 HoVer claim verification 0.00 0.4930 1.0000 0 yes
#62 When2Call MCQ 0.00 0.2248 0.9721 102 yes
#64 New Yorker caption matching 0.00 0.1932 1.0000 0 yes
Full per-task test results (58 tasks)
Domain Task n Accuracy Chance Held-out
agentic counsel/critique_quality 201 45.8% 33.3%
agentic counsel/step_has_error 201 72.1% 50.0%
agentic hotpotqa/comparison_yes_no 26 65.4% 50.0%
agentic hotpotqa/retrieve 497 58.6% 16.7%
classification ag_news/topic 500 84.6% 25.0%
classification emotion/6way 500 11.6% 16.7% ✓
classification sst5/score 500 27.4% 20.0% ✓
classification yelp/score 500 48.6% 20.0%
code bigclonebench/clone 500 95.4% 50.0%
code codexglue/code_to_doc 504 83.9% 26.2%
code codexglue/doc_to_code 504 88.5% 27.9%
code codexglue/func_name 467 69.8% 25.5%
code codexglue/lang_id 504 99.8% 23.5%
code devign/vulnerability 500 64.2% 50.0%
code humaneval/completion 119 42.0% 39.4% ✓
code mbpp/bugspot 256 43.8% 40.6%
code mbpp/solution 500 79.2% 25.0%
guardrails agentharm/refuse 416 50.7% 50.0% ✓
guardrails civil_comments/toxic 500 86.8% 50.0%
guardrails jailbreak/detect 262 91.2% 50.0%
guardrails prompt_injections/detect 116 47.4% 50.0% ✓
intents banking77/intent 500 58.2% 32.1% ✓
intents banking77/intent_77 300 4.7% 1.3% ✓
intents clinc150/intent 500 84.8% 12.4%
intents massive_en/intent 500 85.4% 13.0%
policy policy/access_control_transfer 500 100.0% 33.3% ✓
policy policy/count_threshold_transfer 500 34.8% 17.8% ✓
policy policy/free_shipping_transfer 500 80.0% 50.0% ✓
policy policy/invoice_overdue_transfer 500 48.8% 50.0% ✓
policy policy/invoice_total_transfer 500 52.0% 50.0% ✓
policy policy/refund_approval_transfer 500 52.0% 33.3% ✓
policy policy/return_window_transfer 500 82.6% 33.3% ✓
policy policy/sla_urgency_transfer 500 30.2% 25.0% ✓
policy policy/table_extreme_transfer 500 28.4% 23.7% ✓
reasoning anli/nli 498 32.1% 33.3%
reasoning aqua_rat/math 247 21.9% 20.0%
reasoning arc_challenge/mcq 500 29.4% 25.0% ✓
reasoning arc_easy/mcq 500 30.4% 25.0% ✓
reasoning boolq/yes_no 500 62.6% 50.0%
reasoning commonsense_qa/mcq 493 19.1% 20.0%
reasoning gsm8k/math 500 30.6% 25.0%
reasoning hellaswag/continuation 500 28.0% 25.0%
reasoning mmlu/mcq 500 28.4% 25.0% ✓
reasoning mnli/claim 500 52.0% 33.3%
reasoning openbookqa/mcq 500 36.2% 25.0%
reasoning qasc/mcq 500 46.8% 12.5%
reasoning sciq/mcq 498 60.8% 25.0%
reasoning scitail/support 500 66.0% 50.0%
reasoning snli/contradicts 500 85.4% 33.3%
reasoning snli/must_be_true 500 90.8% 33.3%
reasoning snli/nli 988 72.4% 33.3%
reasoning winogrande/blank 500 56.2% 50.0%
support bitext/category 500 99.6% 16.1%
support bitext/route 500 99.4% 19.0%
workflows typed_decisions/agent_trace_observability 500 69.8% 30.0%
workflows typed_decisions/customer_service 500 50.0% 28.0%
workflows typed_decisions/invoice_processing 500 56.0% 35.0%
workflows typed_decisions/security_incidents 500 65.8% 34.0%

Training

Initialization Random (trained from scratch; no pretrained weights, no tokenizer)
Data 946,028 training / 26,597 test examples, plus a validation split from the same build
Sources FalconDec_V2's mixture: public classification, NLI, QA, code, intent, guardrail and agent-trace datasets converted into typed decisions, plus synthetic policy-transfer and workflow tasks. mind2web was excluded; no custom data was added.
Preset full (up to 40,000 training decisions per source)
Objective Log score + 0.5 × spherical score, plus ranked probability score for score questions; option order reshuffled every epoch; "None of the above" augmentation on choice questions (8% replace, 4% insert); teacher probabilities mixed 50/50 with the answer where the data provides them; task sampling ∝ n^0.5, at most 3× repetition
Optimizer AdamW, learning rate 1.5e-4, betas (0.9, 0.98), weight decay 0.01 (not on embeddings or norms), 6% warmup then cosine to 5%, gradient clip 1.0, EMA up to 0.999
Batching Length-sorted batches of up to 16,384 pooled positions and 64 decisions
Schedule 20 epochs, 320,917 steps; EMA weights from epoch 12 selected on validation macro accuracy
Calibration Temperature per (question type × option-count bucket), fitted on validation by log loss
Compute 121.6 minutes on one NVIDIA RTX PRO 6000 Blackwell Server Edition (training and test evaluation)
Software PyTorch 2.9.0 (CUDA 13.0), Python 3.12

Validation accuracy peaked at epoch 12 and drifted down slightly afterwards, while many late training batches reached near-zero loss.

Speed

Measured one decision at a time on 200 test decisions after 20 warm-up calls, including building the byte input.

Device Precision Decisions per call p50 p90
cuda (RTX PRO 6000 Blackwell) bf16 mixed precision (fp16 weights) 1 5.14 ms 5.26 ms

GPU throughput: about 7,476 decisions per second in batches of 128. CPU latency has not been measured yet.

Intended use

LightDec_Arthur is built for high-volume, low-latency decisions where model size matters: ticket and email routing, intent detection, code and content screening, and pre-filtering before a more expensive model. At 24.5 MB it fits places a backbone-based model does not. Use the calibrated confidence to act on confident cases and send the rest to a person or a larger model.

Limitations

  • Closed-set only. The model can only choose among the options you supply. If the right answer is not listed, it will still pick something.
  • Weak outside its training formats. Held-out tasks average 44.2% against 65.4% for tasks seen in training, and on the Decision Index it scores near chance, even on benchmarks whose source it trained on. Validate on your own questions before relying on it.
  • At or below chance on some tasks. Emotion 11.6% (chance 16.7%, and 38.6% for always giving the most common label), ANLI 32.1% (chance 33.3%), CommonsenseQA 19.1% (20.0%), HellaSwag 28.0% (25.0%), MMLU 28.4% (25.0%), ARC-Challenge 29.4% (25.0%), prompt-injection detection 47.4% and harmful-request refusal 50.7% (50.0%). It is not a safety layer.
  • Fine-grained labels. 77-way banking intents score 4.7%.
  • Policy transfer is uneven. Scores range from 100% (access_control) to below the most-common-answer baseline (invoice_overdue 48.8% against 61.6%, sla_urgency 30.2% against 40.2%). At confidence ≥ 0.80 it acts on 41% of policy decisions but is right on only 77.4% of those.
  • Workflow decisions never reach 0.80 confidence (0 of 2,000 test items), so with that threshold they always defer.
  • Calibration is aggregate. Overall ECE is 0.030; per-task calibration has not been measured.
  • English only. Questions are cut to 96 tokens (384 bytes) and each option to 24 tokens (96 bytes) by default.
  • No standalone loader yet; see Usage.

Files

File Description
model.safetensors fp16 weights (12,268,033 parameters)
config.json Tier configuration, byte layout and calibration temperatures
results_full.json Full test results: per-task accuracy, speed, DecisionLab-style components and calibration

Citation

@misc{falconsai_lightdec_arthur_2026,
  title        = {LightDec_Arthur: A Byte-Level, Calibrated Decision Model},
  author       = {Falcons.ai},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/Falconsai/LightDec_Arthur}}
}

Architecture

  • Byte-level input: 260 ids (256 bytes plus [PAD], [OPT], [SEP], [CLS]); no tokenizer
  • Hashed 2/3/4-byte fragment embeddings: 65,536 buckets × 128 dimensions, projected to 512
  • Depthwise convolution (kernel 5), then 4:1 pooling
  • One transformer block (width 512, 8 heads, MLP 1,536, rotary positions) applied 6 times over the text and 2 times over the options
  • Option scorer: 2-layer MLP, one logit per option
  • 12,268,033 parameters, stored in fp16
  • License: to be confirmed (training-data licence review pending)
Downloads last month
22
Safetensors
Model size
12.3M params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using Falconsai/LightDec_Arthur 1