What Are Positional Encodings in Transformers
SkillVeris Team
AI Research Team

Positional encodings are signals added to token embeddings so a transformer knows the order of words, because self-attention by itself treats input as an unordered set.
In this guide, you'll learn:
- Without them, 'the dog bit the man' and 'the man bit the dog' would look identical to the model.
- The original transformer used fixed sinusoidal functions of different frequencies to encode each position.
- Learned positional embeddings let the model discover its own position vectors during training, at the cost of a fixed maximum length.
- Rotary Position Embedding (RoPE) rotates query and key vectors by an angle tied to position, encoding relative distance directly inside attention.
1What Are Positional Encodings?
Positional encodings are extra numerical signals added to each token's embedding so a transformer can tell which word came first, second, or last. Self-attention computes relationships between all tokens simultaneously and has no built-in sense of order, so without positional information it would treat a sentence as a bag of words.
Concretely, for a token at position p the model produces a vector of the same size as the embedding and combines the two, usually by addition. The result is an embedding that carries both meaning ('cat') and place ('the third token'). Everything downstream then has the ordering it needs to reason about grammar and structure.
2Why Order Matters in Attention
Attention is permutation-invariant: shuffle the input tokens and the raw attention math produces a correspondingly shuffled output with the same values. That is a problem for language, where order carries meaning.
- 'The dog bit the man' and 'The man bit the dog' contain identical tokens but opposite meanings.
- Code, math, and time series all depend on sequence — a for-loop before a return is not the same as after it.
- Recurrent networks handled order naturally by processing one step at a time; transformers trade that for parallelism and must add order back in.
- Positional encodings restore the ordering signal without giving up the speed of parallel attention.
🔑Core Idea
Self-attention sees a set, not a sequence. Positional encodings are how you hand the model the sequence back.
3Sinusoidal (Fixed) Encodings
The original 2017 transformer used fixed sinusoidal encodings: for each position, a vector is built from sine and cosine functions at a range of frequencies. Low-frequency components change slowly across positions, high-frequency ones change quickly, so together they give every position a unique fingerprint.
Because the values are computed from a formula rather than learned, they require no parameters and can, in principle, be evaluated for positions longer than anything seen in training. A useful property is that the encoding for position p+k is a linear function of the encoding for p, which gives the model an easy handle on relative offsets.
- pe(pos, 2i) = sin(pos / 10000^(2i/d))
- pe(pos, 2i+1) = cos(pos / 10000^(2i/d))
- d = embedding dimension, i = dimension index, pos = token position
- No trainable parameters — the pattern is deterministic
4Learned Positional Embeddings
Learned positional embeddings replace the fixed formula with a trainable lookup table: one vector per position, updated by gradient descent like any other weight. Early BERT and GPT models used this approach, and it often fits the training distribution slightly better because the model discovers whatever positional structure the data rewards.
The Trade-off
The catch is a hard maximum length. If you train a table with 512 slots, position 513 simply has no vector, so the model cannot process longer sequences without adding and retraining new rows.
When It Fits
Learned tables are a reasonable default when your inputs have a known, bounded length — classification over short documents, for example — and you do not need to extrapolate.
5Relative and Rotary Encodings
Absolute schemes encode 'this is position 5'. Relative schemes encode 'this token is three steps to the left of that one', which is often what grammar actually depends on. Rotary Position Embedding (RoPE) is the most widely adopted relative method today and appears in many modern open models.
- RoPE rotates the query and key vectors by an angle proportional to their position before the dot product.
- Because rotation composes, the attention score between two tokens ends up depending on their relative distance, not their absolute indices.
- ALiBi takes a simpler route, adding a distance-based penalty to attention scores so far-apart tokens attend less.
- Both extrapolate to longer contexts more gracefully than a fixed learned table.
💡Why RoPE Won
Encoding relative distance inside the attention dot product means the model reasons about 'how far apart' directly — the property most useful for long-context generalisation.
6How Encodings Combine With Embeddings
For absolute schemes, the flow is simple: look up the token embedding, compute or look up the positional vector, and add them element-wise before the first attention layer. The sum flows through the network as a single vector, and attention learns to disentangle content from position as needed.
Rotary encodings work differently — instead of being added to the input, the rotation is applied to queries and keys inside every attention layer. That is why RoPE affects the score computation directly rather than the embedding itself, and why it can be injected at each layer rather than once at the bottom.
7Common Mistakes to Avoid
A few misunderstandings trip up people implementing or fine-tuning models with positional encodings.
- Assuming a learned-position model can handle longer inputs than it was trained on — it usually degrades sharply past its table size.
- Forgetting to apply RoPE consistently to both queries and keys; applying it to only one breaks the relative-distance property.
- Scaling context length without adjusting the RoPE frequency base, which can hurt quality — position interpolation exists for exactly this reason.
- Treating positional encodings as optional — remove them and a transformer's language ability collapses.
- Mixing absolute and relative schemes carelessly, which can send conflicting order signals.
⚠️Watch Out
Extending a model's context window is not just a config flag. The positional scheme has to support the longer range, often through interpolation or retraining.
8Key Takeaways
The essentials of positional encodings come down to a few durable points.
- Self-attention is order-blind; positional encodings supply the sequence information it lacks.
- Sinusoidal encodings are fixed and parameter-free; learned embeddings are trainable but length-capped.
- Relative and rotary schemes encode distance between tokens and extrapolate better to long contexts.
- RoPE rotates queries and keys inside attention and dominates modern long-context models.
- Choice of encoding directly shapes how far a model can reliably read.
9Frequently Asked Questions
Q: Why can't transformers just learn order on their own? A: Self-attention computes all pairwise relationships in parallel with no notion of sequence, so mathematically it treats input as an unordered set. Order has to be injected explicitly through positional encodings, otherwise reordered tokens produce reordered but otherwise identical outputs.
Q: What is the difference between absolute and relative positional encoding? A: Absolute encoding tags each token with its index in the sequence, while relative encoding represents the distance between pairs of tokens. Relative schemes like RoPE and ALiBi tend to generalise better to sequence lengths not seen during training.
Q: Do positional encodings add many parameters? A: Sinusoidal and rotary encodings add essentially no trainable parameters because they come from formulas. Learned positional embeddings add one vector per position, which is small relative to the rest of a large model.
Q: Can I extend a model's context window just by changing positional settings? A: Sometimes, using techniques like position interpolation for RoPE, but it usually needs care and often some additional fine-tuning. A naive length increase past the trained range typically degrades quality.
Related Reading
Get The Print Version
Download a PDF of this article for offline reading.
About the Publisher
SkillVeris Team
AI Research Team
Our AI team covers the latest in machine learning, generative AI, and emerging tech — clearly and accurately.
View all postsRelated Posts
Never miss an update
Get the latest tutorials and guides delivered to your inbox.
No spam. Unsubscribe anytime.