Apache Parquet vs Stroom

Overview

Apache Parquet

Stacks99

Followers190

Votes0

Stroom

Stacks1

Followers3

Votes0

GitHub Stars452

Forks62

Apache Parquet vs Stroom: What are the differences?

What is Apache Parquet? *A free and open-source column-oriented data storage format *. It is a columnar storage format available to any project in the Hadoop ecosystem, regardless of the choice of data processing framework, data model or programming language.

What is Stroom? A scalable data storage, processing and analysis platform. It is a data processing, storage and analysis platform. It is scalable - just add more CPUs / servers for greater throughput. It is suitable for processing high volume data such as system logs, to provide valuable insights into IT performance and usage.

Apache Parquet and Stroom are primarily classified as "Databases" and "Big Data" tools respectively.

Some of the features offered by Apache Parquet are:

Columnar storage format
Type-specific encoding
Pig integration

On the other hand, Stroom provides the following key features:

Receive and store large volumes of data such as native format logs. Ingested data is always available in its raw form
Create sequences of XSL and text operations, in order to normalise or export data in any format. It is possible to enrich data using lookups and reference data
Easily add new data formats and debug the transformations if they don't work as expected

Apache Parquet and Stroom are both open source tools. It seems that Apache Parquet with 992 GitHub stars and 885 forks on GitHub has more adoption than Stroom with 294 GitHub stars and 32 GitHub forks.

Share your Stack

Help developers discover the tools you use. Get visibility for your team's tech choices and contribute to the community's knowledge.

View Docs

CLI (Node.js)

Manual

Detailed Comparison

Apache Parquet	Stroom
It is a columnar storage format available to any project in the Hadoop ecosystem, regardless of the choice of data processing framework, data model or programming language.	It is a data processing, storage and analysis platform. It is scalable - just add more CPUs / servers for greater throughput. It is suitable for processing high volume data such as system logs, to provide valuable insights into IT performance and usage.
Columnar storage format;Type-specific encoding; Pig integration; Cascading integration; Crunch integration; Apache Arrow integration; Apache Scrooge integration;Adaptive dictionary encoding; Predicate pushdown; Column stats	Receive and store large volumes of data such as native format logs. Ingested data is always available in its raw form; Create sequences of XSL and text operations, in order to normalise or export data in any format. It is possible to enrich data using lookups and reference data; Easily add new data formats and debug the transformations if they don't work as expected; Create multiple indexes with different retention periods. These can be sharded across your cluster; Run queries against your indexes or statistics and view the results within custom visualisations; Record counts or values of items over time
Statistics
GitHub Stars -	GitHub Stars 452
GitHub Forks -	GitHub Forks 62
Stacks 99	Stacks 1
Followers 190	Followers 3
Votes 0	Votes 0
Integrations
Hadoop Java Apache Impala Apache Thrift Apache Hive Pig	NGINX MariaDB MySQL IntelliJ IDEA

What are some alternatives to Apache Parquet, Stroom?

MongoDB

MongoDB stores data in JSON-like documents that can vary in structure, offering a dynamic, flexible schema. MongoDB was also designed for high availability and scalability, with built-in replication and auto-sharding.

MySQL

The MySQL software delivers a very fast, multi-threaded, multi-user, and robust SQL (Structured Query Language) database server. MySQL Server is intended for mission-critical, heavy-load production systems as well as for embedding into mass-deployed software.

PostgreSQL

PostgreSQL is an advanced object-relational database management system that supports an extended subset of the SQL standard, including transactions, foreign keys, subqueries, triggers, user-defined types and functions.

Microsoft SQL Server

Microsoft® SQL Server is a database management and analysis system for e-commerce, line-of-business, and data warehousing solutions.

SQLite

SQLite is an embedded SQL database engine. Unlike most other SQL databases, SQLite does not have a separate server process. SQLite reads and writes directly to ordinary disk files. A complete SQL database with multiple tables, indices, triggers, and views, is contained in a single disk file.

Cassandra

Partitioning means that Cassandra can distribute your data across multiple machines in an application-transparent matter. Cassandra will automatically repartition as machines are added and removed from the cluster. Row store means that like relational databases, Cassandra organizes data by rows and columns. The Cassandra Query Language (CQL) is a close relative of SQL.

Memcached

Memcached is an in-memory key-value store for small chunks of arbitrary data (strings, objects) from results of database calls, API calls, or page rendering.

MariaDB

Started by core members of the original MySQL team, MariaDB actively works with outside developers to deliver the most featureful, stable, and sanely licensed open SQL server in the industry. MariaDB is designed as a drop-in replacement of MySQL(R) with more features, new storage engines, fewer bugs, and better performance.

RethinkDB

RethinkDB is built to store JSON documents, and scale to multiple machines with very little effort. It has a pleasant query language that supports really useful queries like table joins and group by, and is easy to setup and learn.

Papertrail

Papertrail helps detect, resolve, and avoid infrastructure problems using log messages. Papertrail's practicality comes from our own experience as sysadmins, developers, and entrepreneurs.

Related Comparisons

Some of the features offered by Apache Parquet are:

Columnar storage format
Type-specific encoding
Pig integration

On the other hand, Stroom provides the following key features:

Receive and store large volumes of data such as native format logs. Ingested data is always available in its raw form
Create sequences of XSL and text operations, in order to normalise or export data in any format. It is possible to enrich data using lookups and reference data
Easily add new data formats and debug the transformations if they don't work as expected

Apache Parquet vs Stroom