Rexium Forge
Specialised AI that runs where your data is.
Ingot is our compressed base model. It fits in 1.45 GB, runs on a laptop or a small server, and answers without sending anything outside your network.
The AI that impresses lives in a datacentre.
It needs a permanent connection, charges for every use, and requires sending your data to someone else's servers. For a clinic, a gym, a factory or anywhere without reliable coverage, that is expensive, risky or simply impossible. Small models that run on cheap hardware do exist, but they are generic: they know a little about everything and little about what your business actually does.
We don't compress everything equally.
A model is not uniform. Some zones act as fundamental memory and take no compression; others are flexible and take plenty. Instead of shrinking everything by the same amount, we measure each zone and compress it exactly as far as it goes. It is the difference between shrinking a house by halving every room and redesigning it knowing which walls are load-bearing.
The result
In our internal exam, Ingot was indistinguishable from a model twice its size while using five times less memory.
The numbers, with the caveat in plain sight
All measured on the same GPU, with the same frozen exam and reference answers from the domain expert.
| Model | Score | Memory |
|---|---|---|
| Qwen3-4B (uncompressed) | 51,7% | 7,50 GB |
| Ingot-2B (6-bit) | 51,7% | 1,44 GB |
| Qwen3.5-2B (uncompressed) | 47,9% | 3,76 GB |
What this does not mean
With a 12-item exam and judge noise measured at about 1.5 points, saying two models tie means we could not separate them, not that they are identical. This is an internal exam, not a public benchmark. We publish it this way because we prefer an honest number to a pretty one, and because it is what any serious buyer will ask.
One base. Many specialists.
Ingot is the raw material, not the finished product. Each business trains on top of it with what it knows, and ends up with a model that understands its domain without ever having shared the data with anyone.
Ingot-2B
The base. Compressed once, the same for everyone, public on request.
Ingot-2B-sport
The first specialist, in progress: high-performance training, from a domain professor's own material, to power the PeakRaptor platform.
Yours
What your business knows and nobody else has. The specialist weights are yours and stay private.
Who this is for
Where the network fails
Basement gyms, factories, construction sites, fieldwork. Places where the connection drops and the work cannot wait for it.
Where data cannot leave
Clinics, health data, data about minors, industrial information. The model travels to the data instead of the data travelling to the model.
Where per-use pricing doesn't add up
High volume, thin margins. A model running on hardware you already own doesn't send an invoice at the end of the month.
Technical sheet
- Base model
- Qwen3.5-2B
- Licence
- Apache-2.0, no revenue cap, no use restrictions
- Format
- GGUF, Q6_K quantisation (6-bit)
- Size
- about 1.45 GB on disk
- Languages
- European Portuguese and English
- Access
- Public on Hugging Face, with manual approval
Ingot derives from Qwen3.5-2B, released under Apache-2.0. What is ours is the compression and the training on top, not the pre-training.
Want a specialist of your own?
Tell us what your business knows and where it needs to run. We reply within 24 hours, and the first reply is from a person.