AbstractPhil
·
AI & ML interests
datasets, research papers, experimentation, vision, classification, text encoders, tokenization, llms, diffusion, distillation, and more.
Recent Activity
posted an update about 5 hours ago After a week of failures and invalid hypothesis with bytelex using generic structures, I found a successful aleph prototypical structure that conforms to the needs. This structure conforms to standard transformer, FFN, and RNN with some minor tweaks.
https://huggingface.co/AbstractPhil/alephllm-mini-beatrix-training/tree/main/mini-beatrix-2s/arms/btx_e003 All the weights of the week are stored here and in various nearby directories.
* We've managed to overlap multiple arms to train multiple simultaneous templates.
* Introduce new tokens as composite tokens from multiple teachers.
* Retrain existing tokens into the behavior of one teacher or another.
* Extend new chains and new behaviors from training in combination.
* Properly instantiate and reinforce behavior using Aleph RNN to reinforce training from raw data.
The EMA Relay. The code has been pushed to both beatrix repos.
https://huggingface.co/AbstractPhil/mini-beatrix-2s/blob/main/relay.py
The structure itself is built specifically as a solidification unit to extensible arms, allowing more composite structures to build.
EMA structures aren't new, but when applied correctly at just such a methodology, the models begin to behave as though the extension relays are in fact the original model. The chains and behavior form naturally and the substructure begins to conform with the token fragments from much more complex structures like combined token differences of T5, Qwen, and CLIP as unified teachers.
The cross-token noise is mitigated using a series of principles and the blueprints are showing both solidity and failure simultaneously, both proving many new utilizable states and disproving multiple theoretical pathologies utilized in current running modern papers as the methodologies tested in the specific formats.
The next article will be fully dedicated to the week of failures leading to the first successes, so stay tuned. repliedto their post about 5 hours ago The post-beatrix-2s and control variant article is finally satisfactory, so the article is now released https://huggingface.co/blog/AbstractPhil/beatrix-ft2
The control variant will need another train with better SDPA stabilization, as the control variant destabilized and collapsed. The primary fault is the lack of QK normalization, which caused the model to simply collapse given enough time. Claude lists the rest of the suspected reasons in the article.
This was a very difficult series of experiments to tune with many fault points. Trying to make heads or tails of Fable 5.1 Claude-speak hasn't been the easiest task either. It seems the model is more likely to create pedantically rigid responses rather than cooperative. Not necessarily insulting, but definitely a sort of refrigerator-magnet behavior - treating my individual contributions as little sketches for the refrigerator. This often completely ignores my larger MD or complex behavioral instructions in favor of my theoretical or hypothetical - likely considering the MD and technical as the model's own, rather than my direct contributions. Right there... right on the refrigerator goes my hypothesis that worked.
https://github.com/AbstractEyes/geolip-bytelex
In any case, this upcoming week will be related entirely to cross-tokenizer distillation research. It may stretch long beyond the next week, but as it stands the geometric vocabulary has evolved into a codebook prediction system.
I would like to give this program linear wings. The Beatrix model supports it, but how well is up for this week to decide.
There are a multitude of potentials based on a series of very recent articles I will be exploring, providing the necessary bytelex complexity to a roughly 60 hour battery of experiments and trainings throughout the geometric systems.
The results will determine the best and worst methodologies of using these models, these shapes, and these structures with more complex byte-level cross tokenization systems View all activity Organizations
view article Twinning Beatrix: A Full-Splat Byte Model, Its Softmax Control, and What Reaches an Image Generator
AbstractPhil
• view article Raising Beatrix: A Byte-Level Model's Measured Childhood
AbstractPhil
• published an article about 1 month ago view article Agreement, Anchors, Addresses: A Week of Geometric Training
published an article about 2 months ago view article Geometric Memory FT4 — Distill Against a Consensus, Ship a Rotation
AbstractPhil
• • 1
published an article about 2 months ago view article The Loss Manifest: A Field History of Objective Functions, and What a Machine Can Actually Be Asked to Compute
published an article about 2 months ago view article Aleph Differentiation, Parts 3 & 3-D: Two Laws, Five Days, One Framework
view article The Aleph Moves Into a Pretrained Trunk: Relays, Registers, and the Two-Regime Dispatch Law
view article The Aleph Under Autoregressive Pressure: Bottleneck Priors, Sign Codes, and the Consumption Law
view article Subject Bucketing: Teaching a Diffusion Model New Prompt Languages Without Forgetting
AbstractPhil
• • 1
view article geolip-aleph-void: The First Relational Geometric Vocabulary Patchwork
view article Reading the Voids: Topological Contribution Signals in Frozen Geometric Codebooks
view article Fused Batched Thin SVD, Part II: Extending the Jacobi Pipeline to N=6 with Configurable Convergence
view article H2 Omega Confirmed, Paradigm Shift: Attempting to Disprove Omega As A Whole
AbstractPhil
• • 1
view article The Polygonal Omega: Trained Sphere-Solvers Are Projective Codebooks
AbstractPhil
• • 1
view article Three Geometric Bands in a Sphere-Normalized Patch Autoencoder
view article The Geometric Engine: Structural Attractors in Neural Network Weight Space
view article FL Hybrid Eigendecomposition Beating cuSOLVER's Mathematical Purity with Compilable PyTorch
view article Ryan Spearman: Geometric Variant Effect Prediction Through Quaternion-Composed Dual Expert Alignment