On 10 December 2025, engineers at Starcloud announced that they had done something no spacecraft was publicly known to have done before: trained a language model in orbit. An NVIDIA H100 aboard the Starcloud-1 satellite learned from a Shakespeare dataset, saved the resulting model and then used it to generate new text while circling Earth.

The processor was not a tiny chip designed only for satellite housekeeping. It belonged to the same H100 family used for AI training and inference in terrestrial data centres. The spacecraft was travelling in low Earth orbit at roughly 500 kilometres altitude, completing a circuit of the planet in about 95 minutes.

The achievement was real, but its scale is easy to misunderstand. This was a compact NanoGPT demonstration, not the orbital equivalent of training ChatGPT. Understanding that distinction makes the experiment more interesting, not less.

Starcloud-1 launched as a 60-kilogram technology demonstrator

Starcloud-1 reached orbit on 2 November 2025 aboard SpaceX’s Bandwagon-4 rideshare mission from Cape Canaveral. The official SpaceX timeline records deployment one hour, 13 minutes and 38 seconds after launch. Nvidia described the spacecraft as weighing about 60 kilograms and being roughly the size of a small refrigerator.

Inside was an H100, a commercial accelerator created for demanding data-centre workloads. Before launch, Nvidia called it the first state-of-the-art, data-centre-class GPU sent into space and estimated that the spacecraft would carry 100 times more GPU compute than any earlier space-based operation.

The satellite is catalogued as Starcloud-1, international designator 2025-248L and NORAD number 66303. Historical orbital data from early December place it at about 507 to 512 kilometres altitude, while tracking data identify a near-circular orbit inclined by roughly 45.4 degrees. Its altitude has gradually declined by a few kilometres since then, as expected for a small spacecraft encountering traces of atmosphere in low Earth orbit.

The model trained in space was NanoGPT, not Gemma

Two different language-model demonstrations are often blended together in accounts of the mission. Starcloud used the H100 to run inference with a preloaded version of Google’s Gemma model. Inference means using weights that have already been learned to produce an answer. The model was operating in space, but its main training had occurred on Earth.

The model trained aboard the spacecraft was Andrej Karpathy’s NanoGPT. Starcloud said it fed the system the complete works of Shakespeare, allowed the model to adjust its weights and then successfully generated Shakespeare-like text from the resulting checkpoint. The company’s mission page now describes Starcloud-1 as the first spacecraft to train a large language model and separately notes the Gemma inference run.

That distinction is the heart of the record. Loading a finished model onto a satellite shows that a spacecraft can execute an AI workload. Training changes the model itself aboard the satellite, exercising repeated forward passes, loss calculation, backpropagation and weight updates. It is a more demanding test of the processor, memory, power system and software stack.

NanoGPT is a teaching-sized transformer, not a frontier model

The phrase “large language model” covers an enormous range. NanoGPT is a clean, hackable implementation of the GPT architecture, but its Shakespeare example is intentionally small. In the public NanoGPT repository, Karpathy calls the default character-level configuration a “baby GPT” and says it can train in about three minutes on one A100 GPU. The input is about one megabyte of Shakespeare text, not a web-scale corpus.

Starcloud has not publicly released enough technical detail to reconstruct its exact run. The company has not published the complete hyperparameter set, model parameter count, training time, energy consumption, loss curve or checkpoint for independent inspection. It is therefore safer to describe the mission as the first publicly reported orbital training of a GPT-style language model than to treat it as a scale record.

This does not make the test trivial. The significance was location. A full training loop ran on a commercial high-performance accelerator hundreds of kilometres from any technician, inside a spacecraft with a strict power budget and no possibility of replacing a failed cable, pump or memory module.

An H100 in orbit faces problems a server rack does not

On Earth, an H100 normally sits inside a heavily engineered server surrounded by high-capacity electrical distribution, liquid or air cooling, redundant networking and technicians. The H100 platform was designed for data centres, where thousands of accelerators can exchange information through specialised interconnects.

A spacecraft replaces that environment with vacuum, radiation and severe limits on mass. Vacuum prevents convective cooling, so waste heat must be conducted away from the GPU and ultimately radiated as infrared energy. Solar panels must produce enough electrical power while batteries and regulators handle changes in illumination and load. Energetic particles can corrupt memory or trigger single-event faults in electronics that were not originally designed as radiation-hardened space hardware.

Starcloud-1 was a useful test because it exposed a real data-centre accelerator to those conditions while running a workload that stressed both computation and memory. A successful short demonstration, however, does not establish a five-year service life. Radiation damage accumulates, thermal cycles repeat and components that work once may still fail under long-duration operation.

Training one small model does not solve orbital scaling

Frontier AI training is a distributed systems problem. Thousands of accelerators must exchange intermediate results at enormous speed, read vast datasets and recover when hardware fails. A single H100 training a small text model avoids most of those networking challenges.

Moving training data into orbit is another constraint. The Shakespeare corpus is small enough to preload or transmit easily. Modern training corpora can contain trillions of tokens, and repeatedly sending large datasets through ground links would consume time, bandwidth and energy. Optical inter-satellite links may eventually connect orbital clusters, but cloud cover, pointing accuracy and ground-station availability still affect the route between space and Earth.

This is why some orbital-computing researchers expect inference or processing of data already collected in space to arrive before large-scale training. Our earlier report on the proposed solar-powered Intelligence Belt reached the same practical divide: answering compact user queries or analysing satellite imagery requires far less data movement than synchronising a frontier training run.

The milestone is best understood as proof of operation

The “first” claim still rests mainly on Starcloud’s announcement, supporting statements from Karpathy and industry reporting. There is no peer-reviewed experiment paper, public telemetry archive or independent reproduction. No comprehensive registry can prove that no earlier classified or unpublished spacecraft experiment trained a small language model. The record should therefore be attributed to Starcloud rather than stated as an independently audited historical certainty.

Even under that boundary, the mission crossed a meaningful line. A commercial accelerator launched on a rideshare flight, entered a roughly 500-kilometre orbit and completed both model training and inference. It showed that the basic AI software stack could survive the transition from a terrestrial server room to an autonomous spacecraft.

What it did not show is equally important. It did not train Gemma from scratch, build a multi-GPU cluster, demonstrate years of reliability or establish that orbital computing is cheaper or cleaner than a ground facility. Starcloud-1 was one refrigerator-sized satellite performing a deliberately manageable job.

That is enough for a first experiment. The history of computing is full of demonstrations whose initial workload looked tiny beside the systems they anticipated. In December 2025, the interesting fact was not that a Shakespeare model became powerful. It was that training crossed the atmosphere and worked.