Skip to content

feat: Add is_same_graph() to compare graphs regardless of vertex and edge order - #2922

Open
schochastics wants to merge 9 commits into
mainfrom
f-349-is-same-graph
Open

schochastics wants to merge 9 commits into
mainfrom
f-349-is-same-graph

Conversation

@schochastics

@schochastics schochastics commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Closes #349.

identical_graphs() compares the internal representation, so it returns FALSE when two graphs store the same vertices and edges in a different order.
Both examples in #349 hit this: graph_from_data_frame() with a reordered vertices argument, and edges added/deleted vs. graph_from_adjacency_matrix().

This PR adds is_same_graph(g1, g2, ..., use_names = TRUE, vertex_attrs = NULL), which checks whether two graphs are the same as labelled graphs:

  • Edge order and endpoint order in undirected graphs are ignored, the number of times each edge occurs is not.
  • Vertices are identified by name if use_names = TRUE and both graphs have a name vertex attribute (via permute()), otherwise by index, as in the C function.
    • Names must be unique, otherwise it errors and points to use_names = FALSE.
    • If only one graph has names, it warns and falls back to indices.
  • vertex_attrs lists vertex attributes that must also be identical after matching up the vertices, e.g. vertex colours.
    • Attributes are compared exactly with identical().
    • An attribute missing from either graph is an error.
    • "name" can be included; this is only meaningful with use_names = FALSE.
  • Graph and edge attributes are not compared. Edge attributes would need a pairing of edges, which is ambiguous with parallel edges.

The docs contrast identical_graphs(), is_same_graph() and isomorphic() as increasingly loose checks, and include an example that canonicalizes both graphs with canonical_permutation() so that is_same_graph() tests for isomorphism.

The C core already has igraph_is_same_graph(), and is_same_graph_impl() was already generated, so there is no C or Stimulus change beyond updating the stale comment in functions-R.yaml.
identical_graphs() stays as it is and now links to is_same_graph() and isomorphic().

Tests are in test-iterators.R and test-aaa-auto.R (for is_same_graph_impl()).

🤖 Generated with Claude Code

schochastics and others added 2 commits September 29, 2026 21:15
…d edge order

`is_same_graph()` checks whether two graphs have the same vertex and edge sets.
Vertices are matched by name when both graphs are named.
It uses the existing `is_same_graph_impl()` from the C core.

Closes #349.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@schochastics
schochastics requested a review from maelle September 29, 2026 19:20
@github-actions

Copy link
Copy Markdown
Contributor

This is how benchmark results would change (along with a 95% confidence interval in relative change) if a228258 is merged into main:

  • ✔️as_adjacency_matrix: 247ms -> 246ms [-1.29%, +0.39%]
  • ✔️as_biadjacency_matrix: 257ms -> 257ms [-0.65%, +0.44%]
  • ✔️as_data_frame_both: 262ms -> 263ms [-0.42%, +1.11%]
  • ✔️as_long_data_frame: 211ms -> 212ms [-0.37%, +1.6%]
  • ✔️es_attr_assign: 240ms -> 238ms [-1.79%, +0.39%]
  • ✔️es_attr_filter: 245ms -> 245ms [-0.45%, +0.34%]
  • ✔️graph_from_adjacency_matrix: 278ms -> 278ms [-0.36%, +0.85%]
  • 🚀graph_from_data_frame: 288ms -> 284ms [-2.39%, -0.27%]
  • ✔️incident_edges: 279ms -> 281ms [-0.83%, +2.58%]
  • ✔️neighbors_by_name: 316ms -> 316ms [-0.93%, +0.76%]
  • ✔️union_named: 238ms -> 235ms [-3.11%, +0.6%]
  • ✔️vs_attr_filter: 329ms -> 330ms [-0.34%, +0.51%]
  • ✔️vs_by_name: 308ms -> 309ms [-0.12%, +0.91%]
    Further explanation regarding interpretation and methodology can be found in the documentation.

@szhorvat

Copy link
Copy Markdown
Member

@schochastics With these AI-generated PRs, do you write the PR description yourself, or does the AI do it?

Open question for reviewers: should graph and vertex attributes be compared too (optionally)? Edge attributes are harder, because pairing edges is ambiguous with parallel edges, so I left them out.

Is this a question from you or the AI?

@maelle

maelle commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

To me a crucial point is the naming. How would user know which one to use, how would they remember? Especially because "same" and "identical" are synonyms in everyday language.

@maelle maelle left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you!!

Comment thread R/iterators.R
.Call(Rx_igraph_identical_graphs, g1, g2, as.logical(attrs))
}

#' Decide if two graphs are the same as labelled graphs

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the title is unclear to me

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The terminology is standard in math and network science. I think this is a good title.

Comment thread R/iterators.R Outdated
#' Decide if two graphs are the same as labelled graphs
#'
#' @description
#' Two graphs are the same if they have the same directedness,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

use an itemized list?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Possible, but my concern is that this would overexplain a standard concept and make it appear as if extra criteria could potentially be listed. In reality the function does something very simple: it compares labelled graphs for equality.

Comment thread R/iterators.R Outdated
#' the same vertex set and the same edge set.
#' Unlike [identical_graphs()], the order in which vertices and edges are
#' stored is ignored,
#' and unlike [isomorphic()], vertices are not relabelled.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

maybe even have a sort of summary table?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

and adding an example case for each, when would you want to use one or the other.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Example would be good. There's an example in the C docs.

Comment thread R/iterators.R Outdated
#'
#' @details
#' If `use_names` is `TRUE` and both graphs have a `name` vertex attribute,
#' vertices are matched by name, so the two graphs may store their vertices

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't understand because this function already ignores the order in which vertices and edges are stored.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some more explanation here would be warranted.

If the vertices are not named, the ith vertex in the 1st graph is considered to be the same as the ith vertex in the second graph. If they are named, and use_names=T, then they are matched up by name.

The quirk is due to the fact that it is possible to have a graph with "unnamed" vertices at all in igraph. Coming from math, this would be a strange concept.

Perhaps a better way to think of it is: vertices may be identified by index, or by name. This flag controlled what to do.

Comment thread R/iterators.R Outdated
#' Vertex names must be unique in both graphs in this case.
#' Otherwise, vertices are matched by their IDs.
#'
#' Edges are compared as a multiset:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is "multiset" a common work in network analysis?

@szhorvat szhorvat Oct 1, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No, that's a computer science term. But what is meant is explained below.

Comment thread R/iterators.R Outdated
#' @return A logical scalar, `TRUE` if the two graphs are the same.
#' @seealso [identical_graphs()] for comparing the internal representation,
#' [isomorphic()] for comparing graphs up to relabelling of the vertices.
#' @cdocs igraph_is_same_graph

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

we don't use the cdocs tag anymore, our roclet adds the link. And if it doesn't it's a roclet bug 😸

Comment thread R/iterators.R
ensure_igraph(g2)
check_bool(use_names)

if (use_names && is_named(g1) && is_named(g2)) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should there be a warning when use_names but they are not both named?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No yes or no from me, just mentioning that in some circumstances vertex indices are auto-converted to string vertex names. So some people may expect to be able to treat the vertex indices of unnamed graphs as vertex "names". Doing such things doesn't make me very comfortable though, so again no "yes" or "no" from me.

Comment thread _pkgdown.yml
- contents:
- graph_id
- identical_graphs
- is_same_graph

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

actually, should we document the three functions (identical, same, isomorphic) on the same page?

@schochastics

Copy link
Copy Markdown
Contributor Author

@schochastics With these AI-generated PRs, do you write the PR description yourself, or does the AI do it?

Open question for reviewers: should graph and vertex attributes be compared too (optionally)? Edge attributes are harder, because pairing edges is ambiguous with parallel edges, so I left them out.

Is this a question from you or the AI?

@szhorvat AI wrote the Description but the question comes from me

@szhorvat

szhorvat commented Oct 1, 2026

Copy link
Copy Markdown
Member

@schochastics With these AI-generated PRs, do you write the PR description yourself, or does the AI do it?

Open question for reviewers: should graph and vertex attributes be compared too (optionally)? Edge attributes are harder, because pairing edges is ambiguous with parallel edges, so I left them out.

Is this a question from you or the AI?

@szhorvat AI wrote the Description but the question comes from me

I just don't want to feel silly answering AI questions assuming they came form you :-)

I won't say a yes or no, but I will try to point out some use cases.

The following would actually make a good example in the documentation of this function (CC @maelle): Take two graphs, canonincalize both (employ canonical_permutation()), and compare with is_same_graph(). This effectively tests for isomorphism.

Now suppose we have vertex coloured graphs (which canonical_permutation() and Bliss support). In this case we would need to compare not only the edge and vertex lists, as is_same_graph() does, but also vertex colours. This is a good use case for optionally allowing attribute comparisons. On the other hand, in some cases one may want to restrict which attributes should be compared.

I hope this helps make the decision.

…s named

Clarify how `use_names` identifies vertices, drop the term "multiset", contrast `identical_graphs()`, `is_same_graph()` and `isomorphic()`, use the example from the C docs and remove the obsolete `@cdocs` tag.
`is_same_graph()` now warns when `use_names = TRUE` but only one graph has vertex names.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@szhorvat

szhorvat commented Oct 1, 2026

Copy link
Copy Markdown
Member

To me a crucial point is the naming. How would user know which one to use, how would they remember? Especially because "same" and "identical" are synonyms in everyday language.

That's indeed rather unpleasant.

Perhaps one way to phrase this is:

  • Isomorphism functions check if two unlabelled graphs are the same. In math we call such pairs of graphs isomorphic.
  • is_same_graphs() checks if two labelled graphs are the same. This is essentially a math concept, but because of the exact issues you asked about regarding use_names, it can't in reality be fully decoupled from the issue of how the graph is represented on the computer. is_same_graphs() will always be false for non-isomorphic graphs, so it's a stricter check
  • identical_graphs() checks if two graphs have the same representation on the computer. This is not math. It is all about the representation, i.e. the data structure. One might say that the naming is misleading as it's not comparing the graphs in the abstract sense, but the igraph data structures. identical_graphs() will always be false for non-is_same-graphs so it's an ever stricter check.

I hope some of the comments were useful but I don't want to barge in. All decisions are yours guys!

schochastics and others added 4 commits October 1, 2026 10:50
`is_same_graph()` gains `vertex_attrs`, a character vector of vertex attributes that must also be identical after matching up the vertices.
Attributes missing from either graph are an error.
The docs add an example that uses `canonical_permutation()` to test for isomorphism, and state that identical, same and isomorphic are increasingly loose checks.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@schochastics

Copy link
Copy Markdown
Contributor Author

(NO AI 🙂)
Ok so I implemented most of what was discussed here, including an added parameter for vertex attributes.

One question:
Should we harmonize the stack of functions to is_identical_graph(), is_same_graph() and is_isomorphic_graph()? Would be a good time to do so for the 3.0.0 release (@maelle, @szhorvat )

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

This is how benchmark results would change (along with a 95% confidence interval in relative change) if df44d42 is merged into main:

  • ✔️as_adjacency_matrix: 183ms -> 181ms [-3.05%, +0.94%]
  • ✔️as_biadjacency_matrix: 198ms -> 199ms [-0.13%, +1.32%]
  • ✔️as_data_frame_both: 202ms -> 203ms [-0.37%, +0.73%]
  • ✔️as_long_data_frame: 161ms -> 164ms [-0.81%, +3.63%]
  • ✔️es_attr_assign: 187ms -> 187ms [-0.79%, +1.35%]
  • ✔️es_attr_filter: 190ms -> 189ms [-0.73%, +0.66%]
  • ✔️graph_from_adjacency_matrix: 219ms -> 218ms [-1.48%, +0.2%]
  • ✔️graph_from_data_frame: 222ms -> 223ms [-0.41%, +1.16%]
  • ✔️incident_edges: 213ms -> 214ms [-0.73%, +1.38%]
  • ✔️neighbors_by_name: 215ms -> 219ms [-0.44%, +3.41%]
  • ✔️union_named: 189ms -> 189ms [-1.69%, +2.16%]
  • ✔️vs_attr_filter: 231ms -> 232ms [0%, +0.97%]
  • ✔️vs_by_name: 212ms -> 211ms [-0.8%, +0.7%]
    Further explanation regarding interpretation and methodology can be found in the documentation.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Identical graphs not considered identical depending on argument vertices in graph_from_dataframe

3 participants