bmarti44 Atlanta, Georgia Consulting Member of Technical Staff, Oracle

Brian
Martin

I build systems that probably should not work.

Frontend architect turned systems and ML engineer. Fifteen years shipping software at scale — eight of them on NBA.com — now spending nights compiling JVMs to WebAssembly and running reinforcement learning on a desktop supercomputer.

Run it where it does not belong
A full JVM inside a browser tab. Doom rendered inside an Oracle database. The interesting part is never the port — it is what breaks when you remove the thing everyone assumed was required.
Frontier models, desktop hardware
Everything in the research column runs on one NVIDIA DGX Spark in my house. No cluster, no credits to burn. Constraints make you measure things.
Results you can argue with
Every experiment measures its own noise floor first, locks the bar before the run, grades against held-out data it never sees, and publishes the nulls. Most of what I try does not work. That is in the write-ups too.

stations 01–04 Selected work

01 / flagship 2026 WebAssembly · OpenJDK · Emscripten · C · TypeScript

JavaBox

A full JVM inside a browser tab.

OpenJDK 21 compiled to WebAssembly. Write Java, compile it, run it — with no server and no JVM installed. Then it runs Doom at 30 frames a second.

cold boot
3–5s
compile + run
<1s
Doom FPS
~30
JDK payload
72MB

Running Java in a browser usually means sending it somewhere else. A server compiles it, runs it, and ships you the output. JavaBox removes the server: it compiles OpenJDK 21’s Zero interpreter to WebAssembly with Emscripten, which means the JVM itself — garbage collector, class loader, bytecode interpreter — executes inside the tab.

The naive version of this is unusable. Booting a JVM per run costs seconds, and nobody waits seconds to see `Hello World`. So JavaBox keeps a persistent CompileServer daemon alive inside the WASM JVM and does compilation in-process through `javax.tools.JavaCompiler`, loading results with a custom in-memory class loader. You pay the boot cost once, then compile-and-run lands under a second.

Then there is the part that makes people check whether it is a video. Mocha Doom — a Doom engine written in pure Java — runs on top of it at roughly 30 FPS, with sound, in a browser, on a JVM that is itself running inside the browser.

JavaBox editor with Java source on the left and program output in a terminal pane on the right
Compile and run, entirely client-side.
Doom, on a JVM, in a tab.

Video

Java renders into a `BufferedImage`. A JNI native method copies the ARGB pixels into the shared WASM heap as RGBA, and the browser draws them with `putImageData` on each `requestAnimationFrame`. No canvas abstraction in the Java layer at all — it thinks it is writing to a framebuffer.

Audio

An 8-channel Java mixer produces 16-bit stereo PCM at 22050Hz. JNI writes it into a shared ring buffer that the Web Audio API drains through an `AudioBufferSourceNode`.

Input, and why there are no threads

Emscripten’s pthread creation is unreliable enough that the usual event-listener-plus-worker design kept deadlocking. Instead the browser pushes key events straight into a C-side queue via a WASM export, and Java polls that queue once per frame through JNI. Single-threaded, boring, and it does not break.

What had to be forked

OpenJDK does not build for Emscripten, so there is a fork of OpenJDK 21u carrying an Emscripten OS abstraction layer, plus a fork of libffi. Both ship as submodules.

Doom, on a JVM, in a tab.

02 / flagship 2026 Oracle MLE · JavaScript · ORDS · TeaVM

DoomDB

Doom, running inside an Oracle database.

Not next to the database. Not storing saves in it. The database is the game engine — it rasterizes every frame and ships the pixels over REST.

FPS over 996 frames
33.2
frames built in-database
320×200
unique frames
995/996
hosting tier
$0

The authoritative Doom world lives in a retained Oracle MLE JavaScript session inside Oracle Autonomous Database 26ai — the Always Free tier. The database advances game state and produces every complete 320×200 indexed framebuffer. ORDS hands back bounded batches of those pixels. The browser’s entire job is applying the palette and blitting to canvas; it does not simulate anything.

A measured run produced 996 frames at 33.159 FPS, 995 of them unique, with every consecutive pair different — which is the bar that separates "it renders" from "it is actually playable."

The engineering is almost entirely latency scheduling. A read-only frame request gets 75ms to finish before it yields to the next Free-tier API lane, which avoids the cancellation backlog that showed up when ORDS requests were aborted mid-flight. The confirmed-pixel ring went from 64 to 128 entries so a slow ORDS tail could not overwrite fresh movement frames while stale lighting frames were still on screen. Retained slots prewarm 96 representative moving and firing frames before signalling READY, so a player’s first turn does not pay for cold portal, weapon, and sprite paths.

Always Free still produces occasional multi-second stalls — one confirmation run recorded 4.594 seconds — and that disqualified the run rather than being smoothed over in the headline. The limitation is in the README because a number you cannot reproduce is not a result.

Doom gameplay rendered in a browser canvas, with frames generated inside Oracle Autonomous Database
Every pixel on screen was computed by a SQL database.

The hosted Always Free instance is gone, so there is no public link to click. The stack runs locally on Docker and Node, and the 33.159 FPS run is re-checkable through the repository’s own acceptance gate.

Every pixel on screen was computed by a SQL database.

03 / flagship 2026 PyTorch · TRL · vLLM · GSPO · Python

Agentic Coder RL

Teaching a 27B model to fix real bugs, graded by real tests.

A from-scratch reinforcement learning pipeline where reward comes from actually running a repository’s test suite in Docker — not from a reward model’s opinion.

Qwen3.6 base
27B
lines of Python
9.1k
test files vs source files
31 / 17
single-GPU target
GH200

Most coding-model training grades output with another model. That is fast and it is also a guess. This pipeline grades by execution: a candidate patch is applied inside a sealed Docker container, the repository’s own tests run, and the reward is whether they pass. It is a slower signal that cannot be talked into agreeing with you.

The stack is an SFT cold start followed by GSPO-token reinforcement learning on SWE-rebench V2, targeting a single GH200-96GB. The interesting constraint is that everything has to fit and stay numerically stable on one GPU, which is where most of the actual work went.

The methodology is the part I would defend in a review. Every phase is gated by a deterministic `make verify-phase-N` script, and an isolated verifier sub-agent re-runs every gate at a different random seed and writes its own report. The main agent is only allowed to advance on an explicit `VERDICT: PASS`. The point is to make it structurally difficult to fool myself about progress.

It is mid-flight: phases 0 and 1 are complete and phase 1.5 — pass-rate calibration — is in progress. There is no trained model to show yet, and I would rather say that than imply otherwise.

Train/inference parity on aarch64

Training in HuggingFace transformers and sampling in vLLM should produce the same logits for the same input. On aarch64 they did not. Finding out why meant bisecting hidden states layer by layer until the divergence boundary appeared, then tracing it into mRoPE and RMSNorm numerics. The whole hunt is checked into `verification/` as roughly 40 JSON and log artifacts, because a debugging session you cannot replay is a story rather than evidence.

Dependency pins as an audit trail

The `pyproject.toml` records why every pin deviates from the original plan. Unsloth was removed entirely because it caps `trl<=0.24` while GSPO-token needs `trl>=1.x`. bitsandbytes went 0.45.1 → 0.46.0 because the aarch64 wheels I was promised did not exist — only x86_64 did.

Teaching a 27B model to fix real bugs, graded by real tests.

04 / flagship 2026 Tauri · Rust · React · TypeScript · Vite

VS Code’s real editor in a native window, with your own model wired in.

Not a Monaco snippet pretending to be an IDE — the actual VS Code workbench running inside a Tauri window, with an AI panel and no Electron.

GitHub stars
35
Electron processes
0
model provider
BYO

Blink runs VS Code’s genuine editor engine through monaco-vscode-api inside a Tauri 2 shell: file explorer, IntelliSense, go-to-definition, multi-cursor, themes, keybindings, and extensions installed from Open VSX. The terminal is a real shell over Tauri’s native PTY. It ships as a macOS `.app`, with Rust rather than a bundled Chromium doing the hosting.

The AI panel sits in the auxiliary bar and is deliberately bring-your-own-key — Anthropic, OpenAI, or anything OpenAI-compatible including a local Ollama endpoint. No proxy in the middle, no subscription attached to the editor.

It is my most-starred project and it is also, by its own README, a buggy proof of concept. Both things are true and I would rather the README say so than have someone discover it at install time.

Blink IDE showing file explorer, code editor, integrated terminal, and an AI chat panel
The whole workbench, plus a chat panel.

station 05 — one DGX Spark, no cluster Research

  1. 01 2026 Python · RL · DGX Spark · agents

    arlab

    An AI research lab that fits on a desk and refuses to flatter itself.

    Give it a hypothesis in plain English. It runs dozens of experiments unattended and then tells you honestly whether anything held up.

    GPU
    1
    possible verdicts
    3
  2. 02 2026 Python · KV cache · Qwen3 · interpretability

    Stencil

    Why models forget the instructions you gave them twenty messages ago.

    A measured study of instruction retention in long conversations — including the part where my trained model lost to a rule with no parameters at all.

    conversations tested
    909
    compliance after eviction
    65% → 17%
    recovered by pinning
    → 60.5%
  3. 03 2026 interpretability · Python · probing · Llama 3.3 70B

    Reference-Free Behavioral Discovery

    Finding behaviors hidden in a model with nothing to compare it against.

    An agent that discovers deliberately implanted hidden behaviors in a 70B model without access to the base checkpoint — tested on Anthropic’s AuditBench.

    AUROC (animal welfare)
    0.911
    topics beam-searched
    243
    reference models used
    0
  4. 04 2026 CUDA · llama.cpp · vLLM · Python

    Frontier at Home

    Frontier-scale models on hardware you can actually buy — with receipts.

    Declarative serving profiles keyed by model, backend, and RAM tier, so the same model comes up correctly on a DGX Spark or a 32GB laptop.

  5. 05 2026 Python · agents · multi-provider

    Research Pipeline

    A hypothesis in English, a preregistered paper out the other end.

    Nine specialized sub-agents, five-round reviews, hash-locked preregistration, and a deterministic gate between every stage.

  6. 06 2026 Python · Clingo · ProbLog · reasoning-gym

    PDCET

    Does training on one task actually help on a different one?

    A task-influence atlas built around a pipeline engineered so that I cannot accidentally cheat — the laptop copy physically cannot run training.

    lines of Python
    16k
    task generators
    39

station 06 — 2007 to now Career

  1. 2021—now

    Oracle

    Consulting Member of Technical Staff

    • IC5. Frontend architecture and platform engineering.
    • Promoted from Principal Frontend Engineer (IC4), 2020–2021.
    • Internal product work, so there is nothing public to link.
  2. 2019—2020

    Delta Air Lines

    Lead UI Engineer — Delta.com

    • Led UI engineering on Delta.com, the airline’s flagship booking and service surface.
  3. 2018—2019

    Turner Sports / WarnerMedia

    Principal Architect — NBA.com

    • Architected the next NBA.com: NBA League Pass as a Dockerized Angular app with server-side rendering, PWA delivery, and hot module replacement.
    • Took that work on the Angular Air podcast (ngAir 201, April 2019).
  4. 2013—2018

    Turner Sports

    Web Developer → Senior → Lead Developer — NBA.com

    • Eight years total on NBA.com across every title in the ladder.
    • Built the NBA Draft 2014 publishing platform on Drupal 7 — live feeds and draft boards for web and mobile.
    • Built the NBA League Pass membership portal against a REST API.
  5. 2011—2013

    Turner

    Web Developer — SportsIllustrated.CNN.com

    • Feature work across SI.com properties including Swim Daily and Extra Mustard.
  6. 2007—2011

    Kennesaw State University

    B.S.

Merged upstream

station 07 — write-ups, including the nulls Writing

Substack

Rip it out by the roots

Detecting planted behavioral modifications in a language model without a reference checkpoint, using cross-layer prediction residuals. AUROC 0.80–0.89 against Anthropic’s AuditBench organisms, p < 0.02.