Understanding Aggregation in Cypher
Module overview
AggregationGrouping rows and computing a value over each group. In Cypher the grouping keys are whatever expressions in the same clause are not aggregated. in CypherNeo4j's implementation of GQL, the ISO standard query language for graph databases. It is declarative: you describe the pattern to find, and the database decides how to find it. means that during a query, the data retrieved is collected or counted. The aggregation of values or counts are either returned from the query or are used as input for the next part of a multi-step query. If you view the execution plan for the query, you will see that rows are returned for a step in the query. The rows that are returned in a query step are an aggregation of propertyA named value stored on a node or a relationship. values, nodesA vertex in a graph. In a property graph it can carry labels and properties., relationshipsA named, directed connection between two nodes. Every relationship has a type, a start node and an end node., or paths in the graph. There are some best practices for aggregating data during your query that you can better understand with examples and explanations of how the query works at runtime.
In this module, you will learn how aggregation works at runtime when using:
collect()collect()vs. subqueries- Using
count() - Using patternA graph structure written in Cypher, such as a node joined to another node by a relationship. comprehension
Data model for this course
This course uses the recommendations dataset for all the queries you will be running and writing. This is the same dataset that will be used for the application development courses in GraphAcademy.
Here is the graph data modelThe labels, relationship types and properties chosen to represent a domain.:

You can view the data model in the sandbox to the right by executing this query:
CALL db.schema.visualization()The node labelsA tag on a node that groups it with other nodes of the same kind. A node can carry more than one. for the graph include:
- Person
- Actor
- Director
- Movie
- Genre
- User
The relationships for the graph include:
- ACTED_IN (with an optional role property)
- DIRECTED (with an optional role property)
- RATED (with rating and timestamp properties)
- IN_GENRE
Also notice that the nodes have a number of properties, along with the type of data that will be used for each property. It is important that you understand the property types defined in the data model.
You can view the property types for nodes in the graph by executing this query:
CALL db.schema.nodeTypeProperties()You can view the property types for relationships in the graph by executing this query:
CALL db.schema.relTypeProperties()Resources
During this course, you can refer to: