Making Configuration Automation Reliable
Made incomplete setup runs fail visibly and preserved supported configuration across export and reapply.
Read initiativeMade incomplete setup runs fail visibly and preserved supported configuration across export and reapply.
Read initiativeAligned plans, warehouse mutations, durable state, recovery, and compiler contracts around fail-closed execution.
Read initiative6,061 GitHub contributions in 2026, flowing through a data pipeline.
Upstream
Selected Projects
An open-source engine for running SQL data transformations safely and correctly in analytics pipelines.
An open standard for data lineage—tracking where data comes from, how it changes, and where it goes.
An open catalog for Iceberg tables—the shared directory that helps engines find and manage lakehouse datasets.
A SPARQL graph database and RDF toolkit for storing and querying knowledge graphs.
An open table format for large analytic datasets, so teams can update and query lakehouse data reliably at scale.
A Spark data-quality toolkit that helps teams catch bad or inconsistent data before it reaches dashboards and models.
A lightweight semantic layer that gives AI agents and people a shared, governed definition of their data.
A Python library that loads data from APIs and databases into analytics warehouses with less glue code.
A fast in-memory Datalog rule engine for analytic data processing with RDF and SPARQL compatibility.
A vendor-neutral YAML specification for exchanging semantic metadata—metrics, dimensions, and relationships—across analytics, AI, and BI platforms.