What we trained · 17 Aug 2026
NVIDIA Nemotron 3.5 Lightning 30B-A3B
An open model, fine-tuned with LoRA supervised fine-tuning in NVIDIA NeMo RL. It's the model behind the FAST setting in Ask the rack, but the tuned version is not in Ask the rack.
Our own numbers, with dates, from a small 3D printing production shop in Austin where we test our tools. Live readings say LIVE. Everything else has a date.
How long the gateway in front of our rack takes to answer, read every 20 seconds.
Each draft is a real call to an open model on our rack. Each decision is a made-up person.
The counts restart at midnight, Austin time. Anything still waiting carries over.
A made-up 60-person CNC shop near Waco gets a steady day of events: web inquiries, orders, alarms, late suppliers, invoices, HR questions, at twelve desks. Si drafts each reply on our rack. Made-up staff decide after a short delay.
The model calls · the drafts · the timing · the rules that route each draft to a named person · the model: NVIDIA Nemotron 3.5 Lightning 30B on DGX Spark 3, Austin
The company · the people · the customers · the events · the know-how it was taught
Nothing in the pane is a recording.
Our first fine-tune runs on our rack today, beside the stock model. Here's what we did, what it shows, and what it doesn't.
NVIDIA Nemotron 3.5 Lightning 30B-A3B
An open model, fine-tuned with LoRA supervised fine-tuning in NVIDIA NeMo RL. It's the model behind the FAST setting in Ask the rack, but the tuned version is not in Ask the rack.
1,269 to train on and 141 held back for checking. They were built, read-only, from an industrial manufacturer's own sales and order records, used with permission (not a customer). Our own open models on the rack drafted the answers, mostly the larger 120B. Nobody's teaching was in it; that is what a Build adds. Every step ran on our rack in Austin, and no rented model saw the data.
Two of our own NVIDIA GB10 machines
Linked by a 200G network, about 12 to 13 hours a run. On one machine alone, it ran out of memory in seven tries (15 Aug 2026), so it took two.
Validation loss on the 141 held-back examples: about 1.66 for the stock model, about 0.78 tuned. A rerun with the same settings gave 0.7751; the first run gave 0.7756.
The tuned model predicts the held-back answers far better, word by word (lower loss is better). Our own models drafted those examples, so this mostly shows it picked up the company's format. It doesn't show the answers are right. Loss is not a business result.
In our own side-by-side read (stock against tuned, on held-back examples), the tuned model answered in the company's working format: the next step, who does it, which system. Its writing didn't get worse. That was our read, not a blind test and not a score.
The fine-tune is a small add-on file (3.7 MB). It was trained on the full-precision open model (62 GB), which it doesn't change. On our rack it runs on that model's 4-bit build, beside the stock model. The scores above are from training, not from the 4-bit build.
Next, a different test: a blind test on about 40 real RFQ packages from a working Texas shop, graded by its estimator and published here whatever it says. It tests the quote desk, not this fine-tune. This fine-tune has no test score yet.
Four NVIDIA GB10 nodes, a control server, switches and storage, in Austin, and a relay server we rent in Germany for the public demo. Each part's job as of 29 Sep 2026; every number has its date.
The machines behind Ask the rack and Brazos Precision.
Four nodes, one control server, the network, the storage, and one rented relay.
Trained our first fine-tune, with DGX Spark 4. Aug 2026.
NVIDIA Nemotron 3 Super 120B, the DEEP setting in Ask the rack.
27 tokens/s · One person asking · Sep 2026
About 55 tokens a second in total across four people asking at once.
NVIDIA Nemotron 3.5 Lightning 30B, the FAST setting and Brazos Precision's drafts, with our first fine-tune beside it, and a 12B vision model.
103 tokens/s · The stock 30B, one person asking · Sep 2026
The small models: a 4B text model, document reading, search, and telling speakers apart in recordings. Trained our first fine-tune, with DGX Spark 1.
The gateway every request passes through, the Brazos Precision simulation, the rack's monitoring and our test shop's software.
Read every 20 seconds from our web host, by way of the relay in Germany.
Switches, and a 200G network between the nodes. Our first fine-tune trained across two nodes over it.
The open models we keep on hand, and the rack's nightly backups.
A small server we rent in Germany, since Aug 2026. The public demo passes through it on its way to Austin. The rack dials out to it, and it lets long answers finish.
Four Si nodes, a control server, switches and storage.
553 W · Sep 2026
About $52 a month of electricity at 13 cents a kWh.
One GB10 server, the kind on our parts list, is rated at 240 W max. That's the maker's rating, not our measurement.
In a small 3D printing production shop in Austin where we test our tools, software does more than 50 jobs. About 20 use a Si model; the rest are plain rules, because rules are better where the answer must be the same every time. A few jobs still use rented superintelligence, such as the agent that coordinates the others; those aren't listed, apart from one row that uses a rented model as backup. September 2026.
Our test shop runs looser rules than a customer's Si will: here, plain rules start a print once a person confirms the plate, the vision check can pause one, routine replies the owner chose go out after a recall window, and quotes under $150 go without the owner. A customer's Si never sends anything, sets a price, buys, or starts, stops or moves a machine (the rules).
| Department | What it does | Where a person decides | Measured |
|---|---|---|---|
| Scheduling Plain rules |
Keeps a three-day line of jobs for every machine: customer orders first, stock second. Paid web orders move up every hour. | A person locks a job, clears a machine or adds a product to the repeat list. | Ran beside the old planner with 0 disagreements before taking over on 24 Sep. |
| Job starts Plain rules |
Starts the next job only when the bed is clear, the right material is loaded and the circuit has headroom. Heat-ups are staggered five minutes apart. | A person confirms the build plate is on before any start. Software never overrides a machine's safety check. | 13 s from plate off to next start. Three machines heating at once drew 2,070 W on one plug, hence the stagger. |
| Quality watch Si, our rack |
A vision model checks each print early and mid-run. On a problem it pauses the machine and texts a photo. | A person taps Resume or Stop. | About 12 s per camera frame. |
| Maintenance Plain rules |
Counts machine hours, projects each service cue from the last 14 days, and builds a Sunday service sheet. | A person ticks each item done. It warns and never blocks the work, except when a waste bin is full. | Cues from 10 to 1,000 hours, sorted into now, 7 days and 30 days. |
| Stock and purchasing Plain rules |
Works out days of cover per colour and builds a reorder from the Sunday count, or sooner under 7 days. | A person approves and checks out. It never buys. | Reorder target: 45 days of use plus one spool, from a 28-day window. |
| Customer messages Si, our rack, a rented model as backup |
Reads the inbox every 30 minutes and drafts replies from the order record. A second model checks text-message drafts. | The owner chose which routine replies may go after a recall window. Safety, legal and money replies always wait for a person. | First run, on a rented model before the move to our rack: 12 threads sorted in 1 min 46 s. |
| Quoting and capacity Si, our rack |
Drafts quotes from the catalogue, the cost model and live capacity. | The owner approves quotes over $150, wholesale and schools. | 11 Sep, on a rented model before the move: room for 420 jobs a month, 480 with another machine. People's time, not machine hours, was the limit. |
| Daily planning Si, our rack, 120B model |
At 05:30 it reads the night's numbers and proposes three things to do today. | A person presses Take or Ignore on each one. | 3 proposals a day, graded by what people take. |
| Approvals Plain rules |
Every two minutes it sorts everything waiting against a written rule file. | The owner holds safety, legal and money; they never default. Carts are capped at $150 each and $400 a month, and a person checks out. | Runs every 2 minutes. |
| Rack health Plain rules |
Checks every machine weekly and texts only when something is wrong. An outside check emails if the rack goes silent for 45 minutes. | A person applies every update except routine overnight security patches. Restarts are always a person's call. | First run found about 32 pending security updates on each Si node. |
A machine paused because it couldn't see a build plate. A camera frame seemed to show one, so our software told the machine to ignore its own check. It was the bare bed, and the nozzle damaged the mat. Since then, no software of ours can override a machine's safety check, and only a person can say the plate is on.
One Monday nothing started, because the start checklist included twelve Sunday cleaning chores. Now maintenance shows as a warning and never blocks the work, except when a waste bin is full. Every Si set-up will ship with rules learned this way.
Five serious outages between 10 Aug and 3 Oct 2026. Each one became something every Si set-up will ship with: memory headroom, monitors that check it's actually answering, and, on your own hardware, remote power control.
On 29 Sep 2026 our own software published a post on our test shop's social accounts without a person approving it. It had filed the post under the wrong rule. We fixed that rule the same day.
On 16 Aug 2026 we lost our first fine-tune's trained weights by our own mistake: they were saved inside a container we removed before checking. We ran it again the next night and got the same result. Now no training run is torn down until the weight files are checked on both machines.
How well it does a real department's work, and whether our fine-tune makes answers better, not just closer to a company's format. Our first test is the quote desk: before our first paid Audit we'll run a blind test on about 40 real RFQ packages from a working Texas shop, graded by that shop's estimator, and publish the count, the date, the model and the score here, whatever it says.
of the up to 3.8 million US manufacturing jobs to fill by 2033 replace people who are retiring. One in four manufacturing workers is 55 or older. Across all US jobs, workers aged 55 to 64 have been with their employer a median 9.6 years; those aged 25 to 34, 3.0.
Sources: Deloitte & The Manufacturing Institute, 2024 · BLS 2025 · BLS 2026
of part buyers expect a quote within 24 hours. When the desk waits on one or two people, customers wait too.
Source: 2020 Part Buyer Survey, via Paperless Parts
of generative AI users at the manufacturers Netskope monitors were still on personal accounts in September 2025, and intellectual property made up a third of the data-policy violations it found going into personal apps. Across industries, MIT NANDA found workers at over 90% of the companies it surveyed using personal AI tools for work.
Sources: Netskope, 2025 · MIT NANDA, 2025, copy at mlq.ai
To talk to the founder, write to human@thinkinghumans.com.