Both models receive the same instruction tuning. Scores are accuracies in percent, with OCRBench divided by ten, and bold marks the better model on each benchmark.
Multimodal large language models learn mostly from paired image-text data or annotated video, and raw web video is rarely used to further train an existing language model. We study whether raw video, with no captions and no text loss, can serve as mid-training data for a pretrained language model. Frames are encoded into continuous visual tokens, and the language model learns to predict the next visual token.
We mid-train Qwen3-1.7B on raw clips from YT-Temporal-1B and then apply the same image-text instruction tuning to it and to the model without mid-training, so that the two differ only in mid-training. The mid-trained model scores 2.9 points higher on average across four video benchmarks and 5.1 points higher across ten image benchmarks, spanning perception, document, and chart tasks. Text performance is preserved even though mid-training includes no text, with an average of 48.9 across 14 text benchmarks compared with 48.0 for the model without mid-training.
Analyses across training show that the image and video gains emerge within 30% of training and plateau thereafter, varying by less than 0.5 points. Predicting captions fails to outperform next-visual-token prediction, demonstrating that video mid-training can remain purely self-supervised without the computational overhead or labeling noise of automated captioning.
A vision encoder \(g_\phi\) and a projector \(p_\psi\) map each frame \(x_t\) to \(N\) continuous visual tokens in the hidden dimension \(d\) of the language model, \(z_t=p_\psi(g_\phi(x_t))\in\mathbb{R}^{N\times d}\). Each token grid is serialized in row-major order with a learned newline embedding after every row, and the frames are concatenated in temporal order into one sequence \(V=[v_1,\dots,v_L]\). The language model \(f_\theta\) reads this sequence causally, and a two-layer MLP head \(r_\omega\) predicts the next visual token from each hidden state \(h_i=f_\theta(v_{1:i})\). Training minimizes the cosine distance \(\mathcal{D}\) between the prediction and the next token, with a stop-gradient \(\mathrm{sg}(\cdot)\) on the target,
$$\mathcal{L}=\frac{1}{L-1}\sum_{i=1}^{L-1}\mathcal{D}\big(r_\omega(h_i),\,\mathrm{sg}(v_{i+1})\big).$$Because attention is causal, the same loss trains next-patch prediction within a frame and prediction across frame boundaries. The objective needs no captions, no discrete codebook, and no text loss, and the prediction head is used only during mid-training.
Setup. Following Cambrian-S, a SigLIP SO400M/14 encoder at 384×384 resolution is connected to Qwen3-1.7B by a two-layer MLP projector. Mid-training uses 16-frame clips sampled at 0.2 frames per second from YT-Temporal-1B, a collection of public YouTube videos, which gives 6.99M clips and 64.4B visual tokens. All parameters are trained end to end for one epoch (27,300 steps) on 128 H100 GPUs, with 256 clips per batch and a 16k-token context. The mid-trained model and the model without mid-training then receive the same instruction tuning on LLaVA-OneVision-Data (3.9M images), so the two differ only in mid-training.
Both models receive the same instruction tuning. Scores are accuracies in percent, with OCRBench divided by ten, and bold marks the better model on each benchmark.
| Video benchmarks | General video QA | Temporal reasoning | Long-form egocentric | Average | |
|---|---|---|---|---|---|
| NExT-QA | VideoMME | TempCompass | EgoSchema | ||
| No mid-training | 59.04 | 42.85 | 51.58 | 45.20 | 49.67 |
| Video mid-training (ours) | 62.18 | 45.15 | 53.04 | 50.00 | 52.59 |
| Gain | +3.14 | +2.30 | +1.46 | +4.80 | +2.92 |
| Image benchmarks | Text-rich image understanding | General perception & diagrams | Character recognition | Complex visual reasoning | Average | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| ChartQA | DocVQA | InfographicVQA | TextVQA | AI2D | MMBench | SEED-Bench | OCRBench | MathVista | MMStar | ||
| No mid-training | 43.76 | 32.13 | 22.31 | 45.49 | 66.39 | 76.00 | 63.00 | 38.40 | 58.15 | 41.60 | 48.72 |
| Video mid-training (ours) | 51.96 | 38.57 | 25.28 | 51.05 | 72.51 | 79.46 | 69.10 | 41.60 | 62.59 | 45.93 | 53.80 |
| Gain | +8.20 | +6.44 | +2.97 | +5.56 | +6.12 | +3.46 | +6.10 | +3.20 | +4.44 | +4.33 | +5.08 |
Mid-training improves every video and image benchmark. The largest video gain is on long-form egocentric understanding, where EgoSchema rises from 45.2 to 50.0. The largest image gains are on chart, document, and diagram tasks such as ChartQA (+8.2), DocVQA (+6.4), and AI2D (+6.1), even though mid-training uses no image-text pairs.
Mid-training contains no text, yet after instruction tuning the mid-trained model reaches a text average of 48.9 over 14 benchmarks, compared with 48.0 for the model without mid-training. Every knowledge and commonsense score stays within 1.4 points of the model without mid-training. Both models fall below the original Qwen3-1.7B on GSM8K, MATH500, and BBH. For the model without mid-training, this loss comes from instruction tuning alone, whose data contain only image-text examples, and the mid-trained model loses less on all three benchmarks.
| Text benchmarks | Qwen3-1.7B | No mid-training | Video mid-training (ours) |
|---|---|---|---|
| Commonsense & general knowledge | |||
| OpenBookQA | 25.99 | 26.79 | 26.59 |
| WinoGrande | 60.53 | 61.87 | 61.64 |
| HellaSwag | 45.18 | 45.26 | 45.68 |
| SIQA | 44.03 | 44.18 | 44.49 |
| PIQA | 71.47 | 72.34 | 72.17 |
| ARC-Easy | 71.58 | 73.31 | 73.61 |
| ARC-Challenge | 39.81 | 41.35 | 42.72 |
| BoolQ | 79.46 | 80.59 | 81.54 |
| MMLU | 62.35 | 61.16 | 60.75 |
| Mathematical & complex reasoning | |||
| MATH500 | 45.63 | 27.98 | 29.17 |
| GSM8K | 73.41 | 49.92 | 54.77 |
| BBH | 38.89 | 23.92 | 29.40 |
| GPQA | 18.00 | 19.50 | 17.50 |
| Code generation | |||
| MBPP | 43.06 | 44.25 | 44.44 |
| Average | 51.38 | 48.03 | 48.89 |
Image and video averages rise by 4.9 and 3.1 points within the first 0.3 epochs (19.5B visual tokens) and change by less than 0.5 points over the remaining 0.7 epochs. The text average stays within 0.9 points of the model without mid-training at every checkpoint. The paper reports the trajectory of each benchmark.
We replace the visual target with a text caption of the current frame or of the next frame. Qwen3-VL-30B-A3B captions consecutive frame pairs with a progress-aware prompt, so the two captions of a pair differ where the action has progressed. All models receive the same instruction tuning, and the table reports the average accuracy on each benchmark suite.
| Mid-training objective | Video | Image | Text |
|---|---|---|---|
| No mid-training | 49.67 | 48.72 | 48.03 |
| Current caption | 51.28 | 48.77 | 47.44 |
| Next caption | 52.51 | 53.54 | 47.40 |
| Visual next token (ours) | 52.59 | 53.80 | 48.89 |
Next-caption prediction reaches video and image averages within 0.3 points of visual next-token prediction, but its text average is 1.5 points lower. At a matched training budget, the 0.5-epoch checkpoint of visual next-token prediction (52.37 video, 53.76 image, 48.32 text) stays within 0.3 points of next-caption prediction on video and image and scores 0.9 points higher on text. Current-caption prediction falls 5.0 points behind on images. Video mid-training can therefore remain fully self-supervised, without the computational cost and label noise of automated captioning.
@article{hwang2026midtraining,
author = {Hwang, Jaedong and Shen, Xiaoqian and Chang, Ernie and Zhao, Changsheng and Zhou, Chong
and Suri, Saksham and Qian, Qi and Liu, Zechun and Wu, Lemeng and Wang, Qinsi
and Krishnamoorthi, Raghuraman and Wen, Wei},
title = {Mid-Training Language Models on Raw Video},
journal = {arXiv preprint arXiv:2610.11019},
year = {2026},
}