AWS Glue logo

AWS Glue

Fully managed extract, transform, and load (ETL) service
62
38
+ 1
0

What is AWS Glue?

A fully managed extract, transform, and load (ETL) service that makes it easy for customers to prepare and load their data for analytics.
AWS Glue is a tool in the Big Data Tools category of a tech stack.

Who uses AWS Glue?

Companies
29 companies reportedly use AWS Glue in their tech stacks, including Postmates, www.autotrader.co.uk, and Plista GmbH.

Developers
30 developers on StackShare have stated that they use AWS Glue.

AWS Glue Integrations

MySQL, Amazon S3, Amazon RDS, Microsoft SQL Server, and Oracle are some of the popular tools that integrate with AWS Glue. Here's a list of all 11 tools that integrate with AWS Glue.

Why developers like AWS Glue?

Here’s a list of reasons why companies and developers use AWS Glue
Top Reasons
Be the first to leave a pro
AWS Glue Reviews

Here are some stack decisions, common use cases and reviews by companies and developers who chose AWS Glue in their tech stack.

AWS Glue
AWS Glue
Amazon EMR
Amazon EMR

I use AWS Glue because I thought it was worth all they hype Fall 2018. However, you had to use Python 2.7 with no pandas support, and cold starts lasted as long as 15 minutes. Also, setting up a dev environment for iterative development was near impossible at the time.

It was a terrible experience for me. I recommend using Amazon EMR instead. Even talking with a friend that works at Amazon, they use EMR instead of Glue for internal spark workloads. Just because a company makes something doesn't mean they use that something :/

See more

AWS Glue's Features

  • Easy - AWS Glue automates much of the effort in building, maintaining, and running ETL jobs. AWS Glue crawls your data sources, identifies data formats, and suggests schemas and transformations. AWS Glue automatically generates the code to execute your data transformations and loading processes.
  • Integrated - AWS Glue is integrated across a wide range of AWS services.
  • Serverless - AWS Glue is serverless. There is no infrastructure to provision or manage. AWS Glue handles provisioning, configuration, and scaling of the resources required to run your ETL jobs on a fully managed, scale-out Apache Spark environment. You pay only for the resources used while your jobs are running.
  • Developer Friendly - AWS Glue generates ETL code that is customizable, reusable, and portable, using familiar technology - Scala, Python, and Apache Spark. You can also import custom readers, writers and transformations into your Glue ETL code. Since the code AWS Glue generates is based on open frameworks, there is no lock-in. You can use it anywhere.

AWS Glue Alternatives & Comparisons

What are some alternatives to AWS Glue?
AWS Data Pipeline
AWS Data Pipeline is a web service that provides a simple management system for data-driven workflows. Using AWS Data Pipeline, you define a pipeline composed of the “data sources” that contain your data, the “activities” or business logic such as EMR jobs or SQL queries, and the “schedule” on which your business logic executes. For example, you could define a job that, every hour, runs an Amazon Elastic MapReduce (Amazon EMR)–based analysis on that hour’s Amazon Simple Storage Service (Amazon S3) log data, loads the results into a relational database for future lookup, and then automatically sends you a daily summary email.
Airflow
Use Airflow to author workflows as directed acyclic graphs (DAGs) of tasks. The Airflow scheduler executes your tasks on an array of workers while following the specified dependencies. Rich command lines utilities makes performing complex surgeries on DAGs a snap. The rich user interface makes it easy to visualize pipelines running in production, monitor progress and troubleshoot issues when needed.
Apache Spark
Spark is a fast and general processing engine compatible with Hadoop data. It can run in Hadoop clusters through YARN or Spark's standalone mode, and it can process data in HDFS, HBase, Cassandra, Hive, and any Hadoop InputFormat. It is designed to perform both batch processing (similar to MapReduce) and new workloads like streaming, interactive queries, and machine learning.
Talend
It is an open source software integration platform helps you in effortlessly turning data into business insights. It uses native code generation that lets you run your data pipelines seamlessly across all cloud providers and get optimized performance on all platforms.
Alooma
Get the power of big data in minutes with Alooma and Amazon Redshift. Simply build your pipelines and map your events using Alooma’s friendly mapping interface. Query, analyze, visualize, and predict now.
See all alternatives

AWS Glue's Followers
38 developers follow AWS Glue to keep up with related blogs and decisions.
Huw Ringer
MAXIM BONDARENKO
Larry Kooper
Himansu Sekhar
Justin Dorfman
John Alton
Sam Chaher
Mohamma76685757
JahnKhan
selyaev