TY - GEN
T1 - NPC-TAG
T2 - 25th IEEE International Conference on Data Mining Workshops, ICDMW 2025
AU - Varolgunes, Uras
AU - Du, Mengnan
AU - Yu, Dantong
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Node classification is a vital task in graph-based deep learning with numerous real-world applications. For node classification tasks on graphs containing text features, it is crucial to design an architecture that can efficiently integrate and process both textual information and graph structure. Recent advancements in pre-trained language models and foundational models have had a significant impact on many fields, including graph and social network domains. Numerous attempts have been made to leverage the robust text processing abilities of these pretrained models. Specifically, for text-attributed node classification tasks, previous works utilize pre-trained language models to either process or enrich the existing textual information. However, many of these approaches require significant computational resources. Inspired by recent work on embedding alignment, we introduce a novel and efficient method, called NPC-TAG (Node Prompts for Classification on Text Attributed Graphs), that aligns the node embeddings generated by a GNN with a frozen large language model and thereby integrates textual information for text-attributed node classification. Our end-to-end design uses the LLM model to process the raw text directly and generates predictions in natural language, while simultaneously capturing the structural information provided by networks. We demonstrate the superior performance of NPC-TAG over the standard GNN pipeline on multiple real-world datasets. For example, we improve the accuracy of the top-ranked method RevGAT from 72.58% to 77.04%, GCN from 71.98% to 76.64%, and GraphSAGE from 72.11% to 76.59% on the ogbn-arxiv dataset, while also showing competitive performance against other SOTA TAG node classification methods.
AB - Node classification is a vital task in graph-based deep learning with numerous real-world applications. For node classification tasks on graphs containing text features, it is crucial to design an architecture that can efficiently integrate and process both textual information and graph structure. Recent advancements in pre-trained language models and foundational models have had a significant impact on many fields, including graph and social network domains. Numerous attempts have been made to leverage the robust text processing abilities of these pretrained models. Specifically, for text-attributed node classification tasks, previous works utilize pre-trained language models to either process or enrich the existing textual information. However, many of these approaches require significant computational resources. Inspired by recent work on embedding alignment, we introduce a novel and efficient method, called NPC-TAG (Node Prompts for Classification on Text Attributed Graphs), that aligns the node embeddings generated by a GNN with a frozen large language model and thereby integrates textual information for text-attributed node classification. Our end-to-end design uses the LLM model to process the raw text directly and generates predictions in natural language, while simultaneously capturing the structural information provided by networks. We demonstrate the superior performance of NPC-TAG over the standard GNN pipeline on multiple real-world datasets. For example, we improve the accuracy of the top-ranked method RevGAT from 72.58% to 77.04%, GCN from 71.98% to 76.64%, and GraphSAGE from 72.11% to 76.59% on the ogbn-arxiv dataset, while also showing competitive performance against other SOTA TAG node classification methods.
KW - graph neural networks
KW - large language models
KW - natural language processing
KW - node classification
KW - Text-attributed graphs
UR - https://www.scopus.com/pages/publications/105035395661
UR - https://www.scopus.com/pages/publications/105035395661#tab=citedBy
U2 - 10.1109/ICDMW69685.2025.00119
DO - 10.1109/ICDMW69685.2025.00119
M3 - Conference contribution
AN - SCOPUS:105035395661
T3 - IEEE International Conference on Data Mining Workshops, ICDMW
SP - 1000
EP - 1009
BT - Proceedings - 25th IEEE International Conference on Data Mining Workshops, ICDMW 2025
PB - IEEE Computer Society
Y2 - 12 November 2025 through 15 November 2025
ER -