Tag: data-engineering
Concepts
- Arrow and In-Memory Formats
- As-Of Joins and Time Alignment
- Bad Tick and Outlier Detection
- Bar Construction Methods
- Batch vs Streaming Pipelines
- Bitemporal Data Modeling
- Brand, Merchant and Subsidiary Mapping
- Building an Alternative-Data ML Pipeline
- Caching and Materialising Feature Tables
- Columnar Storage and Parquet
- Corporate Actions and Price Adjustment
- Data Lineage and Provenance
- Data Quality Checks and Validation
- ETL vs ELT Pipelines
- Event Time vs Ingest Time
- Exchange Calendars and Trading Sessions
- Feature Stores
- Idempotency and Exactly-Once Delivery
- Level 1, Level 2 and Level 3 Data
- Mapping a Vendor Dataset to Your Universe
- Market Data Compression
- Market Data Feed Handlers
- Message Queues and Log-Based Streaming
- How an OHLC Bar Is Actually Built
- Onboarding a Bought Dataset Into Production
- Rebuilding The Book From A Message Feed
- Partitioning and File Layout
- Pipeline Orchestration and DAGs
- Point-in-Time Correctness in Feature Pipelines
- Point-in-Time Databases
- Resampling and Frequency Conversion
- Rolling and Expanding Windows
- Schema Evolution and Versioning
- Streaming Feature Computation
- The Trading Day Boundary and Session Timestamps
- Tick Data vs Bar Data
- Time-Series Databases
- Timestamp Conventions and Time Zones
- Trade and Quote (TAQ) Data
- Universe Construction and Delistings
- Vendor Restatements and Data Revisions