Kyberne Lab / EventMix

Bayesian Predictive Event Transformer

Learn from next-token prediction error inside an isolated session while uncertainty-gated local adapters update across four layers and the pretrained causal backbone remains frozen.

8 causal layers
12 attention heads
15.1M slow parameters
8,192 memory event capacity

bayesian predictive event language model

Learn Locally, Predict with Posterior Confidence

Connecting to the continual Transformer service.

generated text

Autoregressive Output

No generation yet.

slow logits + fast / durable memory

-

runtime state

event-transition-service-v1

Inference mode
-
Returned next token
-
Prompt tokens
-
Generated tokens
-
Learned tokens
-
Memory entries
-
Durable entries
-
Bayesian observations
-
Posterior variance
-
Fast weight norm
-
Session state
-
Elapsed
-

self-supervised deployment checkpoint

Offline Slow Weights. Online Local Bayesian Plasticity.

Tokenizer
-
Gradient parameters
-
Model buffers
-
Selected checkpoint tokens
-
Total training exposure
-
Completed steps
-
Best step
-
Raw validation accuracy
-
Validation PPL
-
Architecture gate
-
Continual learning gate
-

bayesian-predictive-event-transformer-v5

Language Backbone with Multi-Layer Local Bayesian Learning

Byte BPE2,048 learned symbols
Slow Transformer8 causal layers, 12 heads, RoPE and SwiGLU
Causal Successor Pathgeneric within-context token binding and sequence reuse
Bayesian Local Adaptersrank-16 full-covariance updates at layers 2, 4, 6 and 8
Fast + Durable Memorysession delta updates and validated long-term promotion
Uncertainty-Gated Decodeslow logits + successor evidence + local posterior corrections
Model dim
384
Transformer layers
8
Attention heads
12
Context
384 tokens
Slow parameters
15,096,197
Plastic layers
2 / 4 / 6 / 8
Bayesian rank
16, full covariance
Memory capacity
8,192 events
Memory key
128 dimensions
Per-session state
5.34 MB

Offline next-token pretraining learns language and routing features without concept labels. Online learning freezes the backbone, updates fast weights and four rank-16 Bayesian posteriors from local prediction error, and suppresses corrections in uncertain directions. Stable evidence can be retained in a user-isolated durable store. This is a small research model, not an LLM-scale general assistant.