Vtome.ru - электронная библиотека

Learning and Operating Presto: Fast, Reliable SQL for Data Analytics and Lakehouses (Final Release)

  • Добавил: literator
  • Дата: 7-07-2026, 02:20
  • Комментариев: 0
Название: Learning and Operating Presto: Fast, Reliable SQL for Data Analytics and Lakehouses (Final Release)
Автор: Angelica Lo Duca, Tim Meehan, Vivek Bharathan, Ying Su
Издательство: O’Reilly Media, Inc.
Год: 2023
Страниц: 194
Язык: английский
Формат: True PDF, True EPUB
Размер: 10.9 MB

The Presto community has mushroomed since its origins at Facebook in 2012. But ramping up this open source distributed SQL query engine can be challenging even for the most experienced engineers. With this practical book, data engineers and architects, platform engineers, cloud engineers, and software engineers will learn how to use Presto operations at your organization to derive insights on datasets wherever they reside.

Authors Angelica Lo Duca, Tim Meehan, Vivek Bharathan, and Ying Su explain what Presto is, where it came from, and how it differs from other data warehousing solutions. You'll discover why Facebook, Uber, Alibaba Cloud, Hewlett Packard Enterprise, IBM, Intel, and many more use Presto and how you can quickly deploy Presto in production.

With this book, you will:
• Learn how to install and configure Presto
• Use Presto with business intelligence tools
• Understand how to connect Presto to a variety of data sources
• Extend Presto for real-time business insight
• Learn how to apply best practices and tuning
• Get troubleshooting tips for logs, error messages, and more
• Explore Presto's architectural concepts and usage patterns
• Understand Presto security and administration

Deploying Presto to meet your team’s warehouse and lakehouse infrastructure needs is not a minor undertaking. For the deployment to be successful, you need to understand the principles of Presto and the tools it provides. We wrote this book to help you get up to speed with Presto’s basic principles so you can successfully deploy Presto at your company, taking advantage of one of the most powerful distributed query engines in the data analytics space today. The book also includes chapters on the ecosystem around Presto and how you can integrate other popular open source projects like Apache Pinot, Apache Hudi, and more to open up even more use cases with Presto. After reading this book, you should be confident and empowered to deploy Presto in your team, and feel confident maintaining it going forward.

Presto is an open source, distributed SQL query engine that supports structured and semi-structured data sources. You can use Presto to query your data directly where it is located, like a data lake, without the need to move the data to another system. Presto runs queries concurrently through a memory-based architecture, making it very fast and scalable. Within the data lake architecture, you can imagine that Presto fits into the governance and metadata layer. Presto executes queries directly in memory. Avoiding the need for writing and reading from disk between stages ultimately speeds up the query execution time.

The Presto coordinator machine analyzes any query written in SQL (supporting the ANSI SQL standard), creates and schedules a query plan on a cluster of Presto worker machines connected to the data lake, and then returns the query results. The query plan may have a number of execution stages, depending on the query. For example, if your query is joining many large tables, it may need multiple stages to execute, aggregating the tables. You can think of those intermediate results as your scratchpad for a long calculus problem.

Who This Book Is For:
This book is for individuals who are building data platforms for their teams. Job titles may include data engineers and architects, platform engineers, cloud engineers, and/or software engineers. They are the ones building and providing the platform that supports a variety of interconnected products. Their responsibilities include making sure all the components can work together as a single, integrated whole; resolving data processing and analytics issues; performing data cleaning, management, transformation, and deduplication; and developing tools and technologies to improve the analytics platform.

Скачать Learning and Operating Presto: Fast, Reliable SQL for Data Analytics and Lakehouses (Final Release)





ОТСУТСТВУЕТ ССЫЛКА/ НЕ РАБОЧАЯ ССЫЛКА ЕСТЬ РЕШЕНИЕ, ПИШЕМ СЮДА!










ПРАВООБЛАДАТЕЛЯМ


СООБЩИТЬ ОБ ОШИБКЕ ИЛИ НЕ РАБОЧЕЙ ССЫЛКЕ



Внимание
Уважаемый посетитель, Вы зашли на сайт как незарегистрированный пользователь.
Мы рекомендуем Вам зарегистрироваться либо войти на сайт под своим именем.
Информация
Посетители, находящиеся в группе Гости, не могут оставлять комментарии к данной публикации.