Tag: transformers
Concepts
- Adapter Modules and Prefix Tuning
- Attention Entropy and Head Collapse
- Attention Head Specialisation and Probing
- Attention Is Not Explanation
- Autoformer and Series Decomposition Attention
- BERT and Masked Language Models
- Chinchilla Compute-Optimal Scaling
- Do Transformers Beat Linear Baselines for Forecasting?
- Encoder-Only, Decoder-Only and Encoder-Decoder Designs
- FlashAttention and IO-Aware Attention
- FT-Transformer for Tabular Data
- GELU and Swish Activations
- Induction Heads and In-Context Learning
- Informer and ProbSparse Attention
- Instruction Tuning and Supervised Fine-Tuning
- The KV Cache and Incremental Decoding
- Large Language Models
- Linear Attention and Kernel Feature Maps
- Linformer Low-Rank Attention
- Multi-Head Attention
- Multi-Query and Grouped-Query Attention
- Parameter-Efficient Fine-Tuning and LoRA
- PatchTST and Patched Time-Series Transformers
- Performer and Random-Feature Attention
- Pretraining Objectives: MLM, CLM and Span Corruption
- QLoRA and Quantized Fine-Tuning
- The Quadratic Cost of Attention
- Queries, Keys and Values
- Reformer and LSH Attention
- RMSNorm and Scale-Only Normalization
- Rotary Position Embeddings (RoPE)
- The Self-Attention Mechanism
- Sentence Transformers and Contrastive Fine-Tuning
- Sparse Attention: Longformer and BigBird
- Speculative Decoding
- The Transformer Architecture
- Transformer Training Instabilities and Loss Spikes
- Transformers for Cross-Sectional Return Prediction
- Vision Transformers and Patch Embeddings