Databricks Data Engineer Associate: Delta Transactions & Medallion Design
Delta Lake transactions and the medallion pattern solve two different problems that become more valuable when they are used together.…
Data engineering, analytics, databases, machine learning, artificial intelligence and AI systems.
139 published articles in this category.
Delta Lake transactions and the medallion pattern solve two different problems that become more valuable when they are used together.…
Within Databricks, incremental data processing is the practice of computing what changed instead of repeatedly rebuilding everything. At small scale,…
Lakeflow Jobs is Databricks' workflow orchestration layer for coordinating notebooks, Python code, SQL, pipelines, and other tasks as dependable production…
PySpark DataFrames are one of the main programming interfaces for transforming structured and semi-structured data on Databricks. They let engineers…
Schema change is normal in data engineering. Sources add fields, rename concepts, change optionality, widen numeric ranges, and occasionally send…
A slow Spark job is rarely fixed by immediately choosing a larger cluster. Performance problems usually come from a specific…
Unity Catalog is the governance layer that connects identity, permissions, lineage, auditing, discovery, and policy across data and AI assets…
Retrieval-augmented generation, or RAG, improves a generative AI application by retrieving relevant enterprise information and placing that information into the…
Generative AI governance must cover more than model access. Production applications depend on prompts, evaluation datasets, vector indexes, tools, functions,…
Generative AI evaluation is difficult because many important qualities are not captured by one accuracy number. A response can be…
Putting a generative AI model behind an endpoint is only the start of production operations. A reliable serving design must…
Enterprise AI applications often need to find information by meaning rather than exact keywords. Vector retrieval represents text, images, or…
Batch and real-time inference solve different operational problems. Batch scoring processes many records on a schedule or as a dataset…
Machine learning development is iterative. Engineers change training data, features, parameters, algorithms, code, and environments, and each choice can affect…
Change data capture is the point where a lakehouse stops behaving like a sequence of full reloads and starts behaving…