How we research compact models

We work on three problems. Each is judged on quality, memory use and how fast the model runs.

Nox starts by researching algorithms that let large AI models run on much less hardware without losing the capabilities that matter.

Tensor-structured compression

We’re testing whether tensor networks and other tools from quantum information can compress trained models further than pruning, quantization and distillation do. People have already shown this is a serious area of research. Our job is to see whether we can improve on it in a meaningful way.

Capability-preserving evaluation

Before compressing anything, we decide which capabilities matter. Then we report where each method loses something, instead of hiding it in an average. It’s easy to make a model smaller. It’s harder to keep it useful.

Low-compute inference

A smaller model only helps if it runs faster on the hardware people have. We time inference on limited GPUs, because that’s where a method has to work.

How we work

We start with one model, measure what existing methods can do, then test ours against them.

  1. Pick the model. Choose one model and one way of evaluating it before any compression work starts.
  2. Measure existing methods. Find out what current methods already achieve on that evaluation.
  3. Test ours against them. Run our methods under the same conditions and on the same kind of hardware.
  4. Publish all of it. Report quality, memory and speed together, including the results that didn’t work.

How we compare methods

We’ll run every method on the same model and the same hardware, and judge them all the same way. We have no results yet, so nothing here is filled in.

Planned comparison of methods on quality, memory and speed. No results yet.
MethodQualityCapabilities we choose before compressing anything.MemoryWeights and runtime memory, against the GPUs people have.SpeedReal inference time on limited GPUs.
PruningExisting methodNot measured yetNot measured yetNot measured yet
QuantizationExisting methodNot measured yetNot measured yetNot measured yet
DistillationExisting methodNot measured yetNot measured yetNot measured yet
NoxOur methodsPendingNot measured yetNot measured yetNot measured yet

Where we are

We’ve just started. We’re choosing the first model and how to evaluate it. There are no results yet. When there are, we’ll post them here, including the ones that didn’t work.

The Nox team