Tokenizers, model architecture, and quantum optimization — with the same rigor whether the result is a win or a null.
200K-vocab Indic-first multilingual tokenizer. 44 languages, validated on FLORES-200 with a 38-test regression suite.
Read technical deep-dive →A second, independent tokenizer. Beats PU-Tok on 22/30 FLORES-200 languages with 100% round-trip accuracy.
Read technical deep-dive →A 341M model that proposes its own structural changes, backed by statistical evidence — and applies none without human approval.
Read technical deep-dive →QAOA vs. four classical solvers on a realistic exam-scheduling Max-Cut problem, on real IBM Quantum hardware. Classical wins decisively at 42 qubits — we say so.
Read technical deep-dive →15.7B-total / 3.97B-active mixture-of-experts model with domain-specialized experts, upcycled from a dense pretrain.
Read technical deep-dive →A small compiled language — records, optional types, and a full type system, compiling through C to native binaries. Renamed from Arya at v0.5.1.
View on GitHub →A statically-typed language for LLM and data-science workloads — tensors, autodiff, and compile-time shape checking as language primitives. Type checker and Core IR done; not yet runnable.
View on GitHub →© 2026 Cybergeon Technologies. All numbers on this site are re-verified against real code before publishing.