English

Newsgptvis

Method for Visualizing Transformer Internals Using Minimal Data and Simplified Models Released

This article is a translation. Read the Japanese original

The internal mechanisms of Transformer-based large language models—particularly the information flow between layers and the operation of attention mechanisms—are considered difficult to grasp due to the sheer volume of numerical data. This article aims to provide an intuitive understanding of these movements through visualization.

The method employs thorough simplification in three areas: training data, tokenization, and model architecture. The training data consists of a minimal set of 94 words specifically focused on the relationship between fruits and tastes, and a sentence testing the association between "chili" and "spicy" is used for verification.

For tokenization, the method adopts simple word splitting via regular expressions rather than subword methods like BPE. The vocabulary is limited to 19 tokens, with each token corresponding directly to a word.

The model is a decoder-only architecture with an extremely small configuration: two layers, two attention heads per layer, and 20-dimensional embeddings. With approximately 10,000 parameters—orders of magnitude smaller than typical LLMs—it becomes possible to track internal computations.

After 10,000 training steps, the model correctly predicted "chili" in the verification sentence. This indicates that the model learned by generalizing semantic associations rather than through simple memorization.

The code and dataset have been released under the MIT license: https://github.com/rti/gptvis


Source: Understanding Transformers Using a Minimal Example (HN 295pt, 25 comments) (HN Search (backfill), 2025-09-04)