Need advice about which tool to choose?Ask the StackShare community!
Get Advice from developers at your company using StackShare Enterprise. Sign up for StackShare Enterprise.
Learn MorePros of Cassandra
Pros of Elasticsearch
Pros of Apache Spark
Pros of Cassandra
- Distributed119
- High performance98
- High availability81
- Easy scalability74
- Replication53
- Reliable26
- Multi datacenter deployments26
- Schema optional10
- OLTP9
- Open source8
- Workload separation (via MDC)2
- Fast1
Pros of Elasticsearch
- Powerful api327
- Great search engine315
- Open source230
- Restful214
- Near real-time search199
- Free97
- Search everything84
- Easy to get started54
- Analytics45
- Distributed26
- Fast search6
- More than a search engine5
- Highly Available3
- Awesome, great tool3
- Great docs3
- Easy to scale3
- Fast2
- Easy setup2
- Great customer support2
- Intuitive API2
- Great piece of software2
- Reliable2
- Potato2
- Nosql DB2
- Document Store2
- Not stable1
- Scalability1
- Open1
- Github1
- Elaticsearch1
- Actively developing1
- Responsive maintainers on GitHub1
- Ecosystem1
- Easy to get hot data1
- Community0
Pros of Apache Spark
- Open-source61
- Fast and Flexible48
- One platform for every big data problem8
- Great for distributed SQL like applications8
- Easy to install and to use6
- Works well for most Datascience usecases3
- Interactive Query2
- Machine learning libratimery, Streaming in real2
- In memory Computation2
Sign up to add or upvote prosMake informed product decisions
Cons of Cassandra
Cons of Elasticsearch
Cons of Apache Spark
Cons of Cassandra
- Reliability of replication3
- Size1
- Updates1
Cons of Elasticsearch
- Resource hungry7
- Diffecult to get started6
- Expensive5
- Hard to keep stable at large scale4
Cons of Apache Spark
- Speed4
Sign up to add or upvote consMake informed product decisions
- No public GitHub repository available -
What is Cassandra?
Partitioning means that Cassandra can distribute your data across multiple machines in an application-transparent matter. Cassandra will automatically repartition as machines are added and removed from the cluster. Row store means that like relational databases, Cassandra organizes data by rows and columns. The Cassandra Query Language (CQL) is a close relative of SQL.
What is Elasticsearch?
Elasticsearch is a distributed, RESTful search and analytics engine capable of storing data and searching it in near real time. Elasticsearch, Kibana, Beats and Logstash are the Elastic Stack (sometimes called the ELK Stack).
What is Apache Spark?
Spark is a fast and general processing engine compatible with Hadoop data. It can run in Hadoop clusters through YARN or Spark's standalone mode, and it can process data in HDFS, HBase, Cassandra, Hive, and any Hadoop InputFormat. It is designed to perform both batch processing (similar to MapReduce) and new workloads like streaming, interactive queries, and machine learning.
Need advice about which tool to choose?Ask the StackShare community!
Jobs that mention Cassandra, Elasticsearch, and Apache Spark as a desired skillset
What companies use Cassandra?
What companies use Elasticsearch?
What companies use Apache Spark?
What companies use Elasticsearch?
What companies use Apache Spark?
Sign up to get full access to all the companiesMake informed product decisions
What tools integrate with Cassandra?
What tools integrate with Elasticsearch?
What tools integrate with Apache Spark?
What tools integrate with Apache Spark?
Sign up to get full access to all the tool integrationsMake informed product decisions
Blog Posts
What are some alternatives to Cassandra, Elasticsearch, and Apache Spark?
HBase
Apache HBase is an open-source, distributed, versioned, column-oriented store modeled after Google' Bigtable: A Distributed Storage System for Structured Data by Chang et al. Just as Bigtable leverages the distributed data storage provided by the Google File System, HBase provides Bigtable-like capabilities on top of Apache Hadoop.
Google Cloud Bigtable
Google Cloud Bigtable offers you a fast, fully managed, massively scalable NoSQL database service that's ideal for web, mobile, and Internet of Things applications requiring terabytes to petabytes of data. Unlike comparable market offerings, Cloud Bigtable doesn't require you to sacrifice speed, scale, or cost efficiency when your applications grow. Cloud Bigtable has been battle-tested at Google for more than 10 years—it's the database driving major applications such as Google Analytics and Gmail.
Hadoop
The Apache Hadoop software library is a framework that allows for the distributed processing of large data sets across clusters of computers using simple programming models. It is designed to scale up from single servers to thousands of machines, each offering local computation and storage.
Redis
Redis is an open source (BSD licensed), in-memory data structure store, used as a database, cache, and message broker. Redis provides data structures such as strings, hashes, lists, sets, sorted sets with range queries, bitmaps, hyperloglogs, geospatial indexes, and streams.
Couchbase
Developed as an alternative to traditionally inflexible SQL databases, the Couchbase NoSQL database is built on an open source foundation and architected to help developers solve real-world problems and meet high scalability demands.