Journal article · Applied Sciences · 2026

A Review of Graph-Theoretic Approaches in Phishing Website Detection

Magdziarz K., Fraszczak D.

Abstract

Graph-based phishing website detection exploits structural relations that are more difficult for attackers to manipulate than page content alone, including URL-token dependencies, DOM structure, hyperlink neighbourhoods, user–URL interactions and hosting infrastructure. Despite the growing use of graph neural networks, belief propagation and other topology-aware methods in this area, existing phishing-detection surveys treat graph-based approaches only as one category among many and do not provide an internal taxonomy of the graph representations used. This review addresses that gap through a PRISMA 2020-based systematic review of graph-theoretic approaches to phishing website detection published between 2005 and August 2026. Searches were conducted in IEEE Xplore, ACM Digital Library, ScienceDirect and Google Scholar, and were extended through reference-list analysis; 136 records were catalogued and 21 primary studies met all six inclusion criteria. The review proposes a two-axis taxonomy that classifies methods independently by the artefact layer modelled by the graph — address, document, inter-page links, semantic relations, user behaviour, and infrastructure — and by graph type: homogeneous, multi-relational or heterogeneous. Reported accuracies range from 91.0% to 99.7%, but these figures derive from different datasets, class balances and evaluation protocols and therefore cannot be treated as directly comparable. Within-study comparisons show that adding graph structure changes accuracy by −6.1 to +3.9 percentage points, indicating that graph modelling is not uniformly beneficial and that its value depends strongly on the represented layer and deployment context. Only nine studies report processing time and only four evaluate adversarial robustness; in the most severe reported case, accuracy drops from 98.73% to 74.68% under FGSM perturbation. The main barrier to progress is the absence of a shared benchmark with fixed temporal and domain-disjoint splits, standard false-positive reporting and robustness requirements. The review concludes with a concrete benchmark specification and identifies cross-layer graph modelling as the most important unresolved research direction.

Keywords

phishing detectiongraph-based methodsgraph neural networkscybersecurityweb securitymachine learningnetwork analysistaxonomy

Discussion

The first version of this review examined 11 papers in detail. Five reviewers sent 74 separate comments, and by the time I had worked through them, that number had grown to 21. The important part is why it changed.

My original search used six queries built around three terms: “graph,” “PageRank,” and “structural analysis.” That looked reasonable for a review of graph-based phishing detection. It also meant that a paper could use the same underlying idea and remain invisible simply because its authors described it with different vocabulary.

HinPhish exposed the problem. The paper had appeared in Applied Sciences in 2021, the same journal to which I submitted the review, but it described a “heterogeneous information network” rather than a graph. My queries did not retrieve it. NetPhish-Mix, published in February 2026, revealed a second problem. Its model connects domain, URL, and DOM nodes, directly contradicting my initial claim that no prior method combined multiple artifact layers.

Adding two citations would not have fixed the review. I expanded the search from 6 to 14 queries and manually ran 8 new queries across four databases on 17 August 2026. The number of identified records increased from 886 to 1,672, while the detailed corpus grew from 11 to 21 primary studies. I then checked every reported number against the full papers rather than their abstracts. I excluded SpecularNet because the ACM-formatted arXiv preprint provided no evidence of peer review. Three papers whose full texts could not be obtained remained marked as unavailable in the PRISMA flow diagram, not as excluded studies.

The revised search changed the results as well as the corpus. The original accuracy range was 95.5–99.7%; after the expansion, it became 91.0–99.7%. Processing time went from 5 of 11 papers reporting it to 9 of 21. Adversarial robustness went from 1 of 11 to 4 of 21. Those proportions are still poor, but the smaller corpus had made the field look more mature than it was.

For someone building or buying a phishing detector, the headline accuracy of 99.7% is not the useful result. The studies use different datasets, class balances, and evaluation protocols, so their scores cannot support a ranking. Within-study comparisons are more informative because they test graph and non-graph variants under the same conditions. In those comparisons, adding graph information changed accuracy by between −6.1 and +3.9 percentage points. The graph sometimes helped and sometimes made the detector worse. In the most severe reported adversarial test, FGSM perturbation reduced accuracy from 98.73% to 74.68%.

Deployment also depends on what the graph represents. URL- and document-level methods need only the address or page being classified. Link and semantic methods require crawling or external search engines, which makes real-time use harder. Behavioural and infrastructure methods depend on operator data, DNS, WHOIS, or certificates that many defenders cannot access. This is why the review classifies methods along two axes: graph type and artifact layer. The second axis exposes the latency, data requirements, and attack surface hidden behind an accuracy score.

The review ultimately taught me a methodological lesson. Searching for the name I used for a technique was not enough. I also had to search for the names chosen by researchers doing equivalent work. Otherwise, a systematic process can produce carefully verified numbers for an incomplete view of the field.

← All publications

A Review of Graph-Theoretic Approaches in Phishing Website Detection | Krystian Magdziarz