Research

Deep infrastructure work, documented honestly

Tokenizers, model architecture, and quantum optimization — with the same rigor whether the result is a win or a null.

PU-Tok

Production

200K-vocab Indic-first multilingual tokenizer. 44 languages, validated on FLORES-200 with a 38-test regression suite.

Read technical deep-dive →

Hyper-Token

Benchmarked

A second, independent tokenizer. Beats PU-Tok on 22/30 FLORES-200 languages with 100% round-trip accuracy.

Read technical deep-dive →

Self-Evolving LLM

Pretraining

A 341M model that proposes its own structural changes, backed by statistical evidence — and applies none without human approval.

Read technical deep-dive →

Quantum Pipeline

Honest result

QAOA vs. four classical solvers on a realistic exam-scheduling Max-Cut problem, on real IBM Quantum hardware. Classical wins decisively at 42 qubits — we say so.

Read technical deep-dive →

10-Expert MoE

Design locked

15.7B-total / 3.97B-active mixture-of-experts model with domain-specialized experts, upcycled from a dense pretrain.

Read technical deep-dive →

Pcee

v0.8.0

A small compiled language — records, optional types, and a full type system, compiling through C to native binaries. Renamed from Arya at v0.5.1.

View on GitHub →

Prism

Phase 2 complete

A statically-typed language for LLM and data-science workloads — tensors, autodiff, and compile-time shape checking as language primitives. Type checker and Core IR done; not yet runnable.

View on GitHub →

© 2026 Cybergeon Technologies. All numbers on this site are re-verified against real code before publishing.