Rexium Forge

Specialised AI that runs where your data is.

Ingot is our compressed base model. It fits in 1.45 GB, runs on a laptop or a small server, and answers without sending anything outside your network.

1.45 GB on diskApache-2.0European Portuguese and English

The AI that impresses lives in a datacentre.

It needs a permanent connection, charges for every use, and requires sending your data to someone else's servers. For a clinic, a gym, a factory or anywhere without reliable coverage, that is expensive, risky or simply impossible. Small models that run on cheap hardware do exist, but they are generic: they know a little about everything and little about what your business actually does.

We don't compress everything equally.

A model is not uniform. Some zones act as fundamental memory and take no compression; others are flexible and take plenty. Instead of shrinking everything by the same amount, we measure each zone and compress it exactly as far as it goes. It is the difference between shrinking a house by halving every room and redesigning it knowing which walls are load-bearing.

The result

In our internal exam, Ingot was indistinguishable from a model twice its size while using five times less memory.

The numbers, with the caveat in plain sight

All measured on the same GPU, with the same frozen exam and reference answers from the domain expert.

ModelScoreMemory
Qwen3-4B (uncompressed)51,7%7,50 GB
Ingot-2B (6-bit)51,7%1,44 GB
Qwen3.5-2B (uncompressed)47,9%3,76 GB

What this does not mean

With a 12-item exam and judge noise measured at about 1.5 points, saying two models tie means we could not separate them, not that they are identical. This is an internal exam, not a public benchmark. We publish it this way because we prefer an honest number to a pretty one, and because it is what any serious buyer will ask.

One base. Many specialists.

Ingot is the raw material, not the finished product. Each business trains on top of it with what it knows, and ends up with a model that understands its domain without ever having shared the data with anyone.

Ingot-2B

The base. Compressed once, the same for everyone, public on request.

Ingot-2B-sport

The first specialist, in progress: high-performance training, from a domain professor's own material, to power the PeakRaptor platform.

Yours

What your business knows and nobody else has. The specialist weights are yours and stay private.

Who this is for

Where the network fails

Basement gyms, factories, construction sites, fieldwork. Places where the connection drops and the work cannot wait for it.

Where data cannot leave

Clinics, health data, data about minors, industrial information. The model travels to the data instead of the data travelling to the model.

Where per-use pricing doesn't add up

High volume, thin margins. A model running on hardware you already own doesn't send an invoice at the end of the month.

Technical sheet

Base model
Qwen3.5-2B
Licence
Apache-2.0, no revenue cap, no use restrictions
Format
GGUF, Q6_K quantisation (6-bit)
Size
about 1.45 GB on disk
Languages
European Portuguese and English
Access
Public on Hugging Face, with manual approval

Ingot derives from Qwen3.5-2B, released under Apache-2.0. What is ours is the compression and the training on top, not the pre-training.

Want a specialist of your own?

Tell us what your business knows and where it needs to run. We reply within 24 hours, and the first reply is from a person.