DuckDB and Hugging Face: Querying Datasets Directly
DuckDB now supports the hf:// protocol, letting developers run SELECT queries directly against datasets hosted on the Hugging Face Hub without downloading them first. The feature, added in DuckDB v0.10.3, builds on the httpfs extension and works with common formats like CSV, JSONL, and Parquet. Users can query data in place, simplifying ML data pipelines.

DuckDB added support for the hf:// protocol in version 0.10.3, released on May 22 2024. The change builds on the existing httpfs extension and lets a SELECT statement address files that live in a Hugging Face repository without first downloading them. The integration works with common formats such as CSV, JSONL and Parquet, and can even resolve a whole directory of files through a glob pattern. Reading data directly from the Hugging Face Hub removes the need for a separate download step or an external datasets library, simplifying machine‑learning pipelines. Because DuckDB can infer file formats and scan Parquet columns lazily, queries such as counting rows across three files return 173 rows in seconds and can filter on a text pattern to produce 21 matches. The ability to pin queries to a specific branch (for example the ~parquet branch) also gives deterministic access to columnar versions of datasets that were originally stored as CSV or JSONL. It remains unclear how performance will compare with locally stored copies for very large, repeated workloads; the documentation suggests materializing a table for repeated use but does not provide benchmark figures. The handling of private or gated datasets depends on user‑supplied tokens, and no public evidence yet confirms the security implications of storing those tokens in DuckDB’s Secrets Manager.