Open Source Reliability for Data Lake with Apache Spark

Open Source Reliability for Data Lake with Apache Spark

5.962 Lượt nghe
Open Source Reliability for Data Lake with Apache Spark
Open Source Reliability for Data Lake with Apache Spark Presenter: Michael Armbrust of Delta Lake Presented at the Bay Area Apache Spark Meetup hosted at LinkedIn in August 2019. In this talk, they cover: * What data quality problems Delta helps address * How to convert your existing application to Delta Lake * How the Delta Lake transaction protocol works internally * The Delta Lake roadmap for the next few releases Bio: Michael Armbrust is a committer and PMC member of Apache Spark and the original creator of Spark SQL. He currently leads the team at Databricks that designed and built Structured Streaming and the Delta Lake open source project. He received his Ph.D. from UC Berkeley in 2013 and was advised by Michael Franklin, David Patterson, and Armando Fox. His thesis focused on building systems that allow developers to rapidly build scalable interactive applications and specifically defined the notion of scale independence. His interests broadly include distributed systems, large-scale structured storage, and query optimization.