Delta Lake
PulseAugur coverage of Delta Lake — every cluster mentioning Delta Lake across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
Delta Lake Concepts Explained for Databricks Data Engineer Exam
This article serves as a guide to understanding Delta Lake, a crucial component for those preparing for the Databricks Data Engineer Associate Exam. It breaks down the core concepts of Delta Lake, aiming to equip reader…
-
Microsoft Fabric enables enterprise RAG AI with OneLake and Azure OpenAI
A new approach allows for the creation of an enterprise-grade Retrieval-Augmented Generation (RAG) AI system using Microsoft Fabric and OneLake, significantly reducing infrastructure complexity. This method bypasses the…
-
Southern Company enhances storm response with Databricks-powered SCOUT app
Southern Company has developed SCOUT, a real-time storm operations application built on the Databricks Lakehouse platform. This application integrates outage, customer, and reliability data into a single, unified view, …
-
Databricks Project Components Explained: Spark, Delta Lake, MLflow
This article breaks down the core components and technologies that constitute a Databricks project. It highlights the platform's integrated nature, emphasizing tools like Apache Spark, Delta Lake, and MLflow. The explan…
-
Databricks Knowledge Assistant Architecture Enhances Enterprise Search
Databricks has developed a new architecture for its Knowledge Assistant to improve enterprise search capabilities. The system, initially called Instructed Retriever and later refined to Instructed-Retriever-1, addresses…
-
Open Table Formats: Iceberg, Delta Lake, and Hudi Explained
Open table formats like Apache Iceberg, Delta Lake, and Apache Hudi are crucial for data lakes, enabling ACID transactions, schema evolution, and time travel on data stored in object storage. These formats transform col…
-
Data Clustering and Z-Ordering Optimize Table Scans by Enabling Data Skipping
A technical article explains Z-ordering and data clustering, techniques used to optimize data retrieval in large tables, particularly within systems like Delta Lake. It highlights how traditional partitioning can be ins…
-
Amtrak builds AI-powered data backbone for major rail transformation
Amtrak is implementing a comprehensive digital intelligence platform called Rail Intelligence, built on Databricks, to support its most significant operational transformation in over 50 years. This platform aims to unif…
-
Databricks introduces native FILE type for multimodal data in lakehouse
Databricks has introduced a new beta feature called FILE type, designed to store unstructured data like documents, images, and videos as a native column within tables. This innovation aims to integrate multimodal data d…
-
Stop copying Databricks patterns into Microsoft Fabric, advises analysis
Data engineers are advised to avoid replicating complex Databricks architectural patterns within Microsoft Fabric. The article argues that Fabric's unified, serverless, and SaaS-based architecture rewards simplicity ove…
-
Databricks MCP server grants AI agents direct lakehouse access
Databricks has released a new tool, the Databricks MCP server, which allows AI agents like Claude to directly access and interact with a user's data lakehouse. This integration enables conversational execution of notebo…
-
Dow builds carbon ledger on Databricks to track product emissions
Dow has developed a Carbon Footprint Ledger (CFL) utilizing the Databricks Data Intelligence Platform to streamline the calculation of Product Carbon Footprints (PCFs) for its extensive product portfolio. This system, b…
-
Data Versioning Tools for ML and Data Lakes Reviewed
This article provides a practical overview of data versioning tools for machine learning and data lakes. It examines platforms such as Dolt, DVC, Delta Lake, Iceberg, Hudi, and lakeFS, offering insights into their capab…
-
Databricks enables external engine access to Unity Catalog managed tables
Databricks has announced that external engines can now create, read, and write to Unity Catalog managed Delta tables in public preview. This feature, built on open UC Delta APIs, allows engines like Apache Spark, Apache…
-
MLOps CI/CD and Feature Engineering on Azure, AWS, GCP · 2 sources tracked
This cluster explores the implementation of CI/CD pipelines within MLOps across major cloud platforms like Azure, AWS, and GCP. It highlights how MLOps differs from traditional DevOps, emphasizing the importance of feat…
-
Data Lakes vs. Cloud Data Warehouses: Choosing the Right Architecture
This guide compares data lake and cloud data warehouse architectures, highlighting their differences in data storage, query performance, governance, and cost. Data lakes excel at storing raw, multi-format data for machi…
-
Databricks outlines unified data pipeline architecture with Lakeflow
Databricks has detailed its approach to data pipeline architecture, emphasizing a unified platform that integrates batch and streaming data processing. The company highlights that effective architecture separates data i…
-
Databricks launches OpenSharing for AI agents, models, and data
Databricks has introduced OpenSharing, an evolution of its Delta Sharing protocol designed for the agentic AI era. This new open-source protocol, now hosted by the Linux Foundation, expands beyond data sharing to encomp…
-
Text-to-SQL LLM risks: Data leaks and cost overruns
The notion that Text-to-SQL is a solved problem is a dangerous myth, as LLMs can generate non-deterministic SQL queries that pose risks to sensitive data. Approaches like feeding the entire schema to the LLM or using se…
-
Databricks enables unified data access control across engines
Databricks has introduced a beta version of its Cross-Engine ABAC feature, allowing attribute-based access controls to be defined once in Unity Catalog and enforced across various external data engines. This new capabil…