"We don't touch your weights" is a cheap sentence to write and an expensive one for a buyer to check. That applies to us saying it as much as to anyone, which is why the sentence on its own is not worth much.
Here it is with the checkable parts attached. On 2026-08-05 we decided that lab358 would serve stock open-weight checkpoints unmodified — no conversion, no retraining, no fine-tune — and we have held to that since. What follows is what the decision took away from us. That is the part you cannot get from a product page, and it is the part that tells you whether the promise is load-bearing or decorative.
The promise has a date on it, and that is deliberate
What ships is a fact about an artifact, not a property of a company. For anything we produce on or after 2026-08-05, the weights are the upstream weights as published. Before that date, we did convert and retrain, and artifacts from that era may still be in the field.
We could have written the flattering version — lab358 has never modified a weight — and it would have been false. The legal notice that travels with every artifact states both halves instead, per artifact rather than per platform. This is how that statement opens for the SmolLM3-3B family, reflowed from the file's line wrapping and otherwise exactly as it ships, capitals and all:
THIS STATEMENT IS PER ARTIFACT, NOT PER PLATFORM, AND IT IS IN TWO PARTS THAT MUST NOT BE READ AS ONE. FOR AN ARTIFACT PRODUCED ON OR AFTER 2026-08-05 NO WEIGHT IS MODIFIED: the checkpoint's tensors are taken from the pinned upstream revision and loaded unchanged, under their own names, and lab358 applies no retraining, no fine-tuning, no architectural conversion and no quantisation to them. FOR AN ARTIFACT PRODUCED BEFORE THAT DATE the superseded convert-and-retrain lifecycle stated at the end of this paragraph applies instead, and that artifact's converted layers were retrained.
Three things about that quotation, because a post whose subject is overclaiming should not open with a doctored quote.
It is per family, and I have named which one. The paragraph above belongs to the SmolLM3-3B entry. Four of the families in that file carry a before-2026-08-05 history and so state two parts. The other four state it in one part, with no date condition — and the reason is not that we have been serving those stock for longer. It is that no artifact has ever been produced or served through them at all. One family is in service today; the rest of that file is written ahead of the work, so the attribution obligation is discharged the day one of them ships rather than after it.
I stopped quoting before the end of the paragraph, and the rest matters. What follows in the file is the disclosure that the serving process substitutes lab358's own attention path for the checkpoint's stock self-attention module, and that a served artifact carries lab358's retrieval cache — runtime behaviour of our code, altering no stored tensor, and disclosed so that a reader "infers neither that the files were altered nor that nothing differs in how they are computed". Unmodified files and identical computation are two different claims. Only the first one is ours to make.
The wording is not settled. That statement is marked provisional in the file itself, which says it must be re-verified in full against the artifact actually shipped — and reviewed by counsel — before first commercial release, and which records that whether the licence clause it answers "reaches serving-time behaviour at all is an open question for counsel". The per-artifact framing you just read is our current drafting of an open question, not settled instrument text. I would rather publish it with that attached than publish a cleaner sentence without it.
That paragraph is not typed into a document by a person. It is generated from a single authored source in the code, and a test fails when any copy of it drifts from that source. The mechanism exists because we already got this wrong: the same statement was once maintained by hand in three places, the three drifted apart, two of the three ended up false — and the copy that shipped to customers was one of the false ones.
Know what that buys and what it does not. It guarantees that every published copy says what the source says. It cannot guarantee that the source is true of the code; the only automatic check in that direction refuses a short list of retired phrasings by name. The rest is a person reading the implementation, and the file records the date they last did it.
A commitment with a date has a before and an after, which means it can be checked. A commitment without one is just an adjective.
What we deleted
The decision was not a change of emphasis. It removed working code.
A conversion pipeline, a trainer, and a training control plane with its own admin console are all gone from the repository — deleted rather than disabled, because a disabled thing is a thing you intend to switch back on.
One genuine capability went with them. We used to be able to produce a packed artifact: the same model, materially smaller to store and to move. A stock checkpoint arrives in the precision its authors chose, and that is how we ship it. Nothing in the serving layer gives that back, and I am not going to pretend it does.
The remedy we can no longer reach for
This is the cost I would want to know about if I were reading someone else's version of this post.
Before the decision was taken, one of our own engineering findings said, in effect: a model trained one way cannot be served a different way and still be trusted. The fix that finding named was to change the model.
Weeks later we promised not to change any model. So the remedy our own notes prescribed is now permanently unavailable to us, by policy, for every model the platform will ever serve. Not hard — unavailable.
The fact that the finding was written down before the ruling, rather than reconstructed afterwards, is what makes it worth putting in public. It was not a rationalisation. It was a known price, paid on purpose, and it settled a long internal argument in the only direction left open: if we are not allowed to adapt the model to the serving path, the serving path has to adapt to the model. No quiet shortcuts with the model's own computation, because the model never agreed to them.
The bill that arrives every week
Serving a checkpoint unmodified is not a lighter job than modifying one. It is a stricter one. If we do not get to touch the weights, then every piece of arithmetic the model performs has to be reproduced exactly as its authors defined it — by us, by hand, for each model family we support.
We learned what that costs on a rented machine. A large mixture-of-experts checkpoint passed six checks in a row — including one that verifies every tensor in the file by name — and then produced a wrong first token and degenerated into a loop. The seventh check is the one that caught it: a token-for-token comparison of our assembled model against the implementation the checkpoint's own config names. It is the only one of the seven that could have caught this, and it is the one that needed the booking.
The cause was one normalisation step, implemented in the form a neighbouring family uses. Both forms accept a tensor of exactly the same shape, so nothing raised an error. Nothing could have raised an error.
It was also not one mistake but two, which is the part worth taking away. Our first fix corrected the scale term and left the cast order alone. That second half is bit-identical to the correct form at single precision and wrong at the lower precision the fleet actually serves in — by the smallest amount that precision can represent, on every query and every key. A comparison run only at single precision structurally cannot see it. Worse, while the defect was recorded as known we had narrowed the very check that found it down to single precision, and we left it narrowed after the defect was fixed, which took away the check's ability to see the class of thing it was built for. Nothing was red. The coverage was simply gone. We caught that a day later and widened it back, for both precisions, unconditionally.
The response was not a patch. Every block we implement for which we have stated an upstream counterpart is now checked against that family's own published implementation — never against a reference we wrote ourselves, because a reference we wrote inherits whatever we already misunderstood. And both halves of that bug are deliberately re-injected in tests, so a tree that reintroduces either one goes red on a laptop in about the time it takes to import the file.
Three limits on that, stated because a post like this is easy to write without them. The harness runs on synthetic weights on an ordinary CPU: it proves the arithmetic, not the checkpoint, and it is not a substitute for validating a real model on real hardware. For several of the families we have implemented, the check is currently block-by-block rather than end to end. And "for which we have stated an upstream counterpart" is load-bearing: one block in one family has its counterpart named but no comparison behind it yet, and a family that declares no block correspondences at all is checked by nothing in this harness. We would rather write those down than let the sentences before them do more work than they have earned.
We would rather refuse than be subtly wrong
The corollary is a product decision that occasionally annoys people, including us.
If a model arrives in a shape the serving stack cannot reproduce faithfully, it is refused by name. Not a best effort, not a graceful fallback, not a warning in a log that nobody reads.
A subtly wrong model is the worst failure mode in this business, because it does not look like a failure. You get fluent text, at full speed, with a green dashboard, and the only signal that anything is wrong is that the answers are worse than they should be — which you will attribute to the model, because we told you it was the model. A refusal is embarrassing for about a minute. A silent wrong answer is a slow leak of your trust into a place neither of us can see it.
What it costs the pitch
There is a commercial cost too, and it would be strange to write this post and leave it out.
If we do not change your weights, then we do not own anything at the weights layer. If you leave, you keep a model anyone can download. What you would be giving up is the serving stack — which is the honest description of what you are buying, and we would rather compete on that than on the difficulty of getting your own model back.
Including this post
On 2026-09-09, four posts came off this blog. They were deleted rather than edited, including an earlier version of the argument you are reading now.
They described a system the code no longer contained. One of them named a mechanism as live that had been deleted, which is a sentence that was true when it was published and became false while sitting there looking exactly the same. Patching a published post to keep an old claim breathing is how a blog quietly turns into a liability. Deleting it costs traffic and a bit of face, and it is the cheaper of the two.
If this post ages the same way, it should get the same treatment.
What you get for all of that
The reason to accept those costs is that they move some load off your side of the table:
- Quality starts from the model's own record. The published evaluations belong to the checkpoint, and you can run them against it yourself. What they do not settle is how it behaves on our serving path — that is ours to earn, and the comparison that settles it is our serving of a checkpoint against anyone else's serving of the same checkpoint, because the weights are the constant on both sides.
- The weights your review approved are the weights we load. The files are the upstream files, under their own names. What differs is serving-time behaviour — our attention path, and the retrieval cache a served artifact carries — and the notice that travels with the artifact states that in as many words, so your reviewers can read what actually changed instead of taking our word that nothing did.
- A new checkpoint needs engineering, not a training run. When a lab publishes a model you want, what stands between it and your prompt is work on our side: that family's arithmetic reproduced and checked against the family's own implementation, and its licence read. A new family is a project. A further checkpoint inside a family we already serve is smaller — and still not automatic, because provenance is a fact about a checkpoint, not about an architecture.
- Reach is a property of the model you pick — see the model card. What the serving layer adds is durability: index a document once, then retrieve from it at any distance, instead of paying to read it again on every call.
That is the whole trade. We gave up the ability to change your model; what you get back is that the weights are never the variable — and that what our own code does around them is disclosed rather than left for you to assume.
If you have long-document workloads and that trade sounds like the right one, join the lab358 Cloud waitlist — or, to run it self-hosted inside your own AWS account, talk to us.