Benedict Patrick

Small models. Small hardware.

Most AI assumes a data center. These assume a pocket, a breadboard or a browser.

Spark

In training

Plain language in, tool calls out. A tiny model that runs fully offline on an ESP32, built for makers wiring their own devices.

Example

“turn the fan on if it gets hot”

{ "tool": "relay.on", "when": "temp > 30" }

parameters
3.6M
on the chip
2.2 MB
board
$5

Quanta Engine

Running on PC, phone next

An LLM inference engine written from scratch in C++. No llama.cpp underneath. Aimed at the budget Android phones most people actually own.

from scratch
C++
4 bit mixed weights
q4mix
first model, Qwen2.5
0.5B

Relay

In progress

A tool calling model under 100 MB. Grammar constrained decoding means every call it emits actually parses.

target size
<100 MB
parameter class
230M
grammar constrained
JSON

Say hello

Selective about projects. Ambitious about the ones I take: models, engines, and the tools around them.

Start a project →benedictpatrickjohn@gmail.com