Submitted by Remco Hendriks 28 MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes Continker 0 5