Research has been published analyzing the weights of GPT-OSS, the open-weight model released by OpenAI. The researchers demonstrate that details of the training data—described in the model card as "trillions of tokens focusing on STEM, coding, and general knowledge"—can be inferred from the weight parameters.
Analysis of GPT-OSS Weights Reveals OpenAI Training Data Characteristics and Presence of Adult Site Tokens
This article is a translation. Read the Japanese original
In the analysis, the L2 norm distribution of the embedding matrix for the o200k tokenizer used by GPT-5 was examined. It was found that approximately 936 tokens with very low norms were special tokens or specific byte sequences that rarely appeared during the training process. This allows for the estimation of initialization variance and the number of gradient descent steps.
Conversely, tokens with high norms frequently included common English words and terms related to reasoning and coding. This suggests that coding reinforcement learning was the final stage of the training process. Additionally, among non-ASCII characters with high norms, terms from Chinese adult sites and lottery-related content were identified.
These tokens are used commonly across models from GPT-4o onwards. The researchers also demonstrated a phenomenon where presenting specific adversarial inputs results in unintended outputs, as a result of GPT-5 having learned phrases from adult websites.
Source: What GPT-OSS leaks about OpenAI's training data(HN 348pt・82コメント) (HN Search (backfill), 2025-10-06)