PROJECTS

Research systems, inspected in detail.

Each project is presented with its problem, method, evidence boundary, and primary source.

01
Open Model · Decision Systems

Noma:anopen,calibrateddecisionmodel

An open-weight decision model for agents and pipelines. It reads a state and typed questions and returns calibrated probabilities, an abstain signal, and an uncertainty estimate in one forward pass, in about 16 ms, with no generated text.

Decision ModelCalibrationOpen Weights
Noma decision model graphic: state and typed questions pass through a fact channel, an 18-layer backbone and decision heads to a calibrated answer, with median latency against other decision models
Problem

Agents spend most of their calls on small decisions: which queue, which model, did that step work, is it safe to continue. Sending each one to a generative model costs seconds and returns text that has to be parsed and trusted.

Method

Noma keeps the first 18 of 32 layers of a 4B backbone and replaces the language-model head with listwise decision heads, a separate abstain output, and a four-head bootstrap ensemble. A deterministic fact channel supplies dates and quantities, and the state is processed once and shared across questions.

Measured results

16 ms median per decision end to end on one H100; 100% on JevBench easy, 98.6% on the original tier, and 82.6% on a sealed human-reviewed set; $0.023 per 1,000 decisions by JevBench's method.

Evidence boundary

The figures are our own measurements with the JevBench client and are not yet on the leaderboard. Noma is built for single-pass decisions: on multi-step reasoning items it scores 46.4% on a held-out set, and those belong with a reasoning model.

02
Research · Transformers

CausalInformationFlowinTransformerResidualStreams

Causal experiments using activation patching, intervention, probing, and ablation to study how specialist attention heads interact through GPT-2's shared residual stream.

InterpretabilityGPT-2Causal Analysis
Residual-stream causal map showing dominant routed paths, secondary paths, and the direct-path intervention site across transformer layers
Paper overview

A causal study of trajectory-conditioned semantics and dynamic computational pathways in GPT-2 Small.

Methodology

Activation patching, targeted intervention, linear probing, ablation, and topology analysis are used to trace information through the shared residual stream.

Measured results

The public preprint reports its measured results and limitations. The primary paper remains the source of record for every quantitative claim.

Interactive implementation / visualization

AXON visualizes sparse-autoencoder feature activations token by token. It is an implementation artifact of this research, not a separate research claim.

03
Research System · Reasoning

LEMMA

A neuro-symbolic mathematical reasoning system combining transformer-guided proposal generation, Monte Carlo Tree Search, symbolic rules, and verifiable state transitions.

ReasoningMCTSNeuro-Symbolic
LEMMA guided-search tree showing transformer proposals, Monte Carlo Tree Search expansion, symbolic verification, pruned candidates, and verified solutions
Problem

Mathematical reasoning needs broad search while every accepted transformation must remain explicit and verifiable.

Transformer proposals

A learned proposal mechanism prioritizes promising next transformations without replacing symbolic validation.

MCTS and symbolic rules

Monte Carlo Tree Search explores candidate paths over an extensible rule system.

Verification loop

Each state transition is checked symbolically before it can advance the reasoning trace.

04
Open Source · Simulation

GS-DroneGym

Photorealistic aerial-agent simulation combining 6-DOF drone dynamics, Gaussian Splatting rendering, waypoint supervision, and synthetic VLA trajectory generation.

SimulationVLA3DGS
GS-DroneGym research graphic showing behavior-cloning loss, synthetic dataset composition, and the closed-loop aerial-agent environment
Simulator overview

A drone-first environment for embodied-learning research and reproducible trajectory generation.

Dynamics and rendering

Six-degree-of-freedom dynamics run against Gaussian Splatting scenes to narrow the visual gap between simulation and reconstructed environments.

Observation and action flow

Episodes align observations, state, language, waypoints, and safety annotations along a shared trajectory.

Dataset and adapters

The pipeline exports synthetic episodes and supports benchmark-oriented data adapters described in the public repository.

05
Experimental System · Model Behaviour

Activation-spaceinterventionsforfrozenlanguagemodels

An experimental framework for SAE-guided activation intervention and causal-control tests in frozen language models.

ExperimentalSAECausal Control
BrainPatch activation-space intervention experiment diagram showing the frozen-model pipeline, feature 727 and random-control comparisons, and the measured behavior shifts
Experimental framework

The system edits selected activation-space features in a frozen model and measures downstream behaviour.

SAE-based intervention

Sparse-autoencoder features provide intervention targets without updating model parameters.

Causal-control methodology

Controlled comparisons test whether an activation edit produces the intended behavioural transfer.

Negative result

The experiment did not validate reliable behavioural steering. That failure boundary is preserved as part of the evidence.