Instructions to use Codingstark/LFM2.5-350M-PR-Origin with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Codingstark/LFM2.5-350M-PR-Origin with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("LiquidAI/LFM2.5-350M") model = PeftModel.from_pretrained(base_model, "Codingstark/LFM2.5-350M-PR-Origin") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Unsloth Studio
How to use Codingstark/LFM2.5-350M-PR-Origin with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Codingstark/LFM2.5-350M-PR-Origin to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Codingstark/LFM2.5-350M-PR-Origin to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Codingstark/LFM2.5-350M-PR-Origin to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="Codingstark/LFM2.5-350M-PR-Origin", max_seq_length=2048, )
LFM2.5-350M PR Origin Classifier
This is a LoRA adapter for LiquidAI/LFM2.5-350M. It classifies the likely
origin of a pull request as:
codex: OpenAI/Codex family;claude: Anthropic/Claude family;unknown: human, mixed, unsupported, or insufficient evidence.
It supports a Model Inversion routing experiment:
- predicted Codex PR -> Claude reviewer;
- predicted Claude PR -> GPT reviewer;
- unknown or low confidence -> default or human review.
This model estimates authorship style. It does not prove who authored code.
Important limitations
OSAPRD cohort labels are heuristic rather than cryptographically verified. Results may include repository, time, agent-version, and collection biases. Mixed human/AI authorship and adversarially rewritten PRs are particularly difficult. Evaluate on controlled generation logs, later-time data, and unseen repositories before production use.
Liquid AI notes that the base checkpoint is not recommended for general programming. This adapter uses it only for a narrow structured classification task; it is not a code-generation model.
Evaluation
The notebook loaded 500 Claude, 500 Codex, and 500 unknown PRs. Exact deduplication left 1,483 records, split by repository into 890 train, 316 validation, and 277 test examples.
| Metric | Result |
|---|---|
| Raw accuracy | 0.9386 |
| Raw macro-F1 | 0.9376 |
| Selective macro-F1 | 0.9513 |
| Selective coverage | 0.9206 |
| 10-bin ECE | 0.0205 |
| Multiclass Brier score | 0.0827 |
| Character n-gram baseline macro-F1 | 0.8282 |
Selective evaluation uses validation-fitted temperature 0.1278, minimum
probability 0.80, and minimum top-two margin 0.15. See metrics.json and
run_manifest.json for machine-readable details.
These pilot metrics are not evidence of reliable attribution outside the sampled OSAPRD cohorts.
Usage
pip install -r requirements.txt
python inference.py example_pr.json
The CLI loads the public base checkpoint and this adapter, masks explicit model names from the SLM input, scores all three canonical JSON answers, applies the validation-fitted calibration, and abstains when confidence is low.
Direct PEFT loading:
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
adapter_id = "Codingstark/LFM2.5-350M-PR-Origin"
base_id = "LiquidAI/LFM2.5-350M"
tokenizer = AutoTokenizer.from_pretrained(adapter_id)
base = AutoModelForCausalLM.from_pretrained(
base_id,
dtype=torch.float16,
device_map="auto",
)
model = PeftModel.from_pretrained(base, adapter_id)
model.eval()
Training
- Base revision:
b9d6e4e2d75f440b12a2b4d731c808004ecbbd89 - Dataset revision:
1c8ed7b6963ae31e4b601fdbfdbe83b8e8817c82 - Method: Unsloth 16-bit LoRA
- LoRA rank/alpha: 16/16
- Trainable parameters: 5,013,504
- Sequence length: 2,048
- Epochs/steps: 3/336
- Seed: 3407
- Accelerator used: Google Colab Tesla T4
Executed notebook: Google Colab
Dataset: OSAPRD/OSAPRD
Base model: LiquidAI/LFM2.5-350M
License
This derivative is distributed under the LFM Open License v1.0. Review the
included LICENSE, preserve required attribution, and confirm the current
commercial-use conditions before redistribution or deployment. Source
repositories represented in OSAPRD retain their own licenses.
- Downloads last month
- 7