Abstract
-
Purpose
This study examined the knowledge structure and thematic characteristics of health literacy research in Korea through keyword network analysis and topic modeling.
-
Methods
Keyword frequency and co-occurrence network analyses included 637 articles. For latent Dirichlet allocation, English abstracts were combined with standardized author keywords. After preprocessing and document-frequency filtering, 636 articles remained in the effective latent Dirichlet allocation corpus. Models with 3–10 topics were evaluated across 10 random seeds according to log-likelihood, perplexity, semantic coherence, topic distinctiveness, cross-seed stability, and interpretability of representative documents. A five-topic solution was selected, and the final model was estimated using seed 2026. An abstract-only sensitivity analysis was conducted with 633 articles.
-
Results
Health literacy occupied the most central position in the network, as expected because it was the core search concept. Among the remaining keywords, older adults, eHealth literacy, self-efficacy, health behavior, and self-care had the highest degree centrality. Louvain clustering identified six thematic communities. The five latent Dirichlet allocation topics were Health Literacy Measurement, Mental Health Literacy, Health Information and Communication, Digital Health Literacy, and Disease-Specific Health Literacy. Topic prevalence ranged from 17.1% to 21.4%. Disease-Specific Health Literacy was the largest topic (21.4%, n=136), followed by Digital Health Literacy (21.2%, n=135) and Mental Health Literacy (21.1%, n=134). The sensitivity analysis produced a broadly comparable five-topic structure.
-
Conclusion
This study systematically mapped the knowledge structure of Korean health literacy research. The findings shed light on major research themes and their relationships and may help inform future research priorities.
-
Key Words: Health literacy; Korea; Social network analysis; Natural language processing; Health communication
INTRODUCTION
Health literacy is a major determinant of health outcomes and health equity because it affects how individuals access, understand, evaluate, and apply health information when making decisions [
1-
4]. In patients with chronic conditions, health literacy is associated with treatment adherence, self-care, and disease management outcomes [
5-
7]. Limited health literacy has been linked to poor medication adherence, increased hospitalization, and adverse health outcomes, particularly among older adults and other vulnerable populations [
7-
9]. Higher health literacy, by contrast, is associated with better self-management, higher quality of life, and more effective engagement with health care systems [
10,
11].
Health literacy includes basic functional skills but also extends to broader competencies, such as digital health literacy, eHealth literacy, organizational health literacy, and critical health literacy [
12-
14]. This expanded scope reflects the increasing complexity of health care environments and the growing use of digital technologies to obtain and manage health information. As the concept has broadened, numerous studies have synthesized and conceptualized the field through systematic reviews and conceptual analyses [
14-
16].
In Korea, health literacy research has addressed health information comprehension, health behaviors, health care utilization, chronic disease self-management, and health inequities across diverse populations, including older adults, people with chronic diseases, and other vulnerable groups [
17,
18]. Recent studies have also examined digital health literacy, health information accessibility, and strategies for reducing health disparities [
19,
20].
Despite this breadth, few studies have systematically examined the knowledge structure of health literacy research in Korea, including its major thematic domains and the relationships among research topics. Korean studies have reviewed health literacy research trends and policy implications [
18] and have used topic modeling to identify major themes in domestic health literacy research [
21]. However, these studies have not comprehensively examined the intellectual structure of the field or the interrelationships among research themes through an integrated approach combining keyword network analysis and topic modeling.
International studies have used bibliometric analysis, keyword network analysis, and topic modeling to examine the knowledge structure and thematic patterns of health literacy research [
22,
23]. By analyzing large-scale textual data, these approaches can identify major research themes and clarify how they are related [
24,
25]. However, most existing analyses have focused on English-language publications indexed in international databases. Such studies may not adequately capture the health care system, policy context, or sociocultural characteristics that shape Korean health literacy research. Because health literacy is strongly influenced by health care systems and sociocultural environments [
1,
4], findings from international literature may not fully represent the distinctive features of the Korean context.
A comprehensive analysis of literature indexed in the Korean Citation Index (KCI) is therefore needed to characterize the knowledge structure and major thematic domains of health literacy research in Korea. This study used keyword network analysis and topic modeling to examine health literacy-related articles published in KCI-indexed journals. Specifically, it examined structural relationships among keywords, identified major research themes, and characterized the knowledge structure of Korean health literacy research. By mapping the intellectual landscape of this field, the study may help identify research gaps and inform evidence-based directions for future scholarship and policy development.
METHODS
1. Study Design
This text-mining study characterized the current landscape of health literacy research in Korea. Word cloud analysis, keyword network analysis, and topic modeling were used to identify major research themes and latent topic structures.
2. Data Sources and Search Strategy
Relevant literature was retrieved from the KCI, a national database of peer-reviewed academic journals published in Korea. Searches were conducted with the KCI “All Fields” option to identify health literacy-related research in Korea as comprehensively as possible. Four Korean search terms related to health literacy were searched separately: “건강문해력” (health literacy), “건강정보이해” (understanding of health information), “건강정보활용” (use of health information), and “헬스 리터러시” (health literacy). The number of records retrieved for each search term is shown in
Supplementary Figure 1. The initial searches were conducted on April 16, 2026, and an additional search was conducted on June 1, 2026, to capture newly indexed publications. Publication year was not restricted; therefore, all records indexed in KCI from database inception through June 1, 2026, were eligible for screening.
The search identified 1,235 records. Records retrieved from multiple searches were merged into a single dataset, and duplicates were removed using article title, author names, publication year, and journal title. After 401 duplicate records were removed, 834 unique articles remained for screening.
Two researchers independently reviewed titles, abstracts, and author-provided keywords to assess eligibility. Discrepancies in study selection were resolved through discussion and consensus. Studies were included when health literacy was a major focus, including studies of its measurement, determinants, outcomes, interventions, or conceptual development.
Studies were excluded when (1) health literacy was mentioned only incidentally and was not a major focus; (2) relevance to health literacy could not be determined from the title, abstract, or keywords; or (3) the publication was not an original scholarly article, such as an editorial, announcement, conference notice, or other nonacademic material.
After title, abstract, and keyword screening, 197 articles were excluded, yielding a final study dataset of 637 articles. All 637 articles were retained for keyword-based analyses because standardized author-provided keywords were available. Four articles lacked usable English abstracts, but their standardized author keywords were retained for construction of the combined latent Dirichlet allocation (LDA) corpus. After preprocessing and document-frequency filtering, one of these keyword-only records contained no retained terms; therefore, the effective LDA corpus included 636 articles. The overall study selection process is shown in
Supplementary Figure 1.
3. Data Preprocessing
For keyword-based analyses, author-provided keywords were cleaned and standardized in four steps. First, keywords were separated into individual units using comma delimiters. Second, duplicate keywords were removed. Third, synonymous expressions and spelling variants were standardized with a customized thesaurus developed through iterative review. Two researchers independently reviewed the thesaurus, and discrepancies were resolved through discussion and consensus. Examples of standardized variants included “self care” and “self-care,” “eHealth literacy” and “e-health literacy,” and “elderly,” “older people,” and “older adults.” Related Health Literacy Survey instrument terms, including HLS-EU, HLS-Q12, and HLS-SF12, were standardized under the umbrella term “HLS.” Fourth, differences in capitalization and spelling were harmonized across records.
Closely related terms were standardized, but conceptually distinct terms were kept as separate keywords. For example, self-care and self-management were not merged because self-management is considered a distinct dimension of self-care in chronic disease management.
For topic modeling, additional normalization procedures were used to preserve multiword concepts during tokenization. Frequently occurring multiword terms identified during thesaurus development were converted into single lexical units with underscore notation, including “self efficacy” → “self_efficacy,” “quality of life” → “quality_of_life,” and “older adults” → “older_adults.” Standardized author keywords were then merged with the corresponding English abstracts using a unique article identifier. This unified corpus was designed to represent both author-defined research themes and the substantive content of the publications.
The corpus was then preprocessed by lowercasing text, removing punctuation and noninformative characters, and applying a customized stop-word dictionary. The stop-word list was refined during preliminary model development and finalized before the sensitivity and stability analyses. Non-substantive reporting and methodological terms identified during this process—including confirmed, mediation_analysis, t_test, correlated, review, and similar generic descriptors—were removed along with standalone numeric expressions. Terms were retained when they could convey substantive meaning in the health literacy literature. The final preprocessing specification was frozen and applied consistently to the full and abstract-only corpora.
4. Keyword Frequency and Word Cloud Analysis
Keyword frequency analysis was performed using the standardized author-provided keyword dataset to identify the most frequently occurring terms in the health literacy literature. A word cloud was then generated from the frequency distribution to visually summarize the relative prominence of keywords.
Because health literacy was the primary search term used to identify the literature and overwhelmingly dominated the frequency distribution, it was excluded from the word cloud to make secondary research themes more visible. Keywords appearing fewer than two times were excluded, and the 100 most frequent remaining keywords were retained to improve interpretability. The resulting word cloud showed the relative prominence of key concepts within the field.
5. Keyword Network Analysis
Keyword network analysis was performed using standardized author-provided keywords. For each article, unique keywords were identified, and all possible keyword pairs were generated to construct a weighted, undirected document-level co-occurrence network. In this network, nodes represented keywords, edges represented co-occurrence within the same article, and edge weights indicated how often keyword pairs co-occurred across articles.
To improve interpretability and reduce network sparsity, only keyword pairs with a co-occurrence frequency of three or greater were retained. After network construction, nodes with fewer than two connections were removed from the final visualization. Node size was categorized by keyword frequency (<10, 10–19, 20–49, 50–99, and ≥100 occurrences), edge width represented co-occurrence strength, and node color indicated communities identified through Louvain community detection.
Community structure was identified by applying the Louvain community detection algorithm [
26] to the final weighted keyword co-occurrence network, with edge weights representing co-occurrence frequencies. A random seed of 2026 was set before community detection to ensure reproducibility. The modularity statistic was calculated to assess the strength of the resulting community structure. Figure 2 and
Supplementary Table 1 were generated from the same final network object and identical Louvain community membership.
Network centrality measures—including degree centrality, betweenness centrality, closeness centrality, and eigenvector centrality—were calculated from the same final keyword co-occurrence network to identify structurally prominent keywords and examine their roles within the network. Degree centrality was calculated as the unnormalized number of direct connections between a keyword and other keywords. Betweenness centrality reflected the extent to which a keyword bridged different parts of the network, whereas closeness centrality indicated its proximity to all other keywords; both measures were calculated using the normalized options implemented in the igraph package. Eigenvector centrality reflected the relative importance of a keyword based on its connections to other highly connected keywords and was calculated using eigen_centrality() in the igraph package [
27].
6. Topic Modeling
Topic modeling was performed using LDA [
28] to identify latent thematic structures in Korean health literacy research. A document-term matrix was constructed from the preprocessed corpus and used as the input for LDA modeling.
To determine the topic structure, LDA models with 3–10 topics were estimated using Gibbs sampling [
29] implemented in the topicmodels package in R. Candidate solutions were evaluated across 10 random seeds. The burn-in period was 1,000 iterations, the number of sampling iterations was 3,000, and the thinning interval was 100. Dirichlet priors were set to α=50/k and β=.1 [
29]. For the selected five-topic model (k=5), this specification corresponded to a symmetric α parameter of 10 for each topic. Model selection considered log-likelihood, perplexity, semantic coherence of high-probability terms, topic distinctiveness, cross-seed stability based on one-to-one matching of topic-specific top terms, and substantive interpretability. Because statistical fit improved monotonically as k increased, the final solution was not selected on the basis of minimum perplexity alone. Models adjacent to the selected solution were also examined for topic merging and fragmentation. After k=5 was selected, the final model was estimated with a fixed random seed of 2026 for reproducibility.
Topic interpretation was based on the terms with the highest β values and the 10 documents with the highest posterior probabilities (γ) for each topic. Preliminary labels were assigned according to the dominant concepts represented in both sources. Two researchers independently reviewed the labels and representative documents, and discrepancies were resolved through discussion and consensus.
7. Statistical Analysis and Visualization
All analyses were conducted using R software (version 4.5.2, R Foundation for Statistical Computing, Vienna, Austria). Data preprocessing and text-mining procedures were performed using the tidyverse and tidytext packages. Topic modeling was conducted using the topicmodels package. Keyword network analyses, including centrality-measure calculation and community detection, were performed using the igraph package, whereas network visualizations were generated using ggraph. Visualization outputs included word clouds, keyword co-occurrence networks, topic-term plots, and topic-distribution bar charts.
8. Ethical Considerations
This study analyzed publicly available bibliographic data and did not involve human participants. The Institutional Review Board (IRB) of Samsung Medical Center advised that the study was not subject to IRB review because it used only publicly available bibliographic data.
RESULTS
1. Word Cloud for Keywords
A word cloud was generated to summarize the most frequently occurring author-provided keywords across the included studies (
Figure 1). Because health literacy was the most frequent keyword and dominated the visualization, it was excluded from the word cloud to improve the interpretability of the remaining terms. After health literacy was excluded, the most frequent keywords were “older adults,” “eHealth literacy,” “health behavior,” “self-efficacy,” “social support,” and “mental health literacy” (
Table 1). Other frequent terms included “self-care,” “health promotion,” “diabetes mellitus,” “quality of life,” and “nursing students.”
2. Keyword Co-occurrence Network Analysis
Figure 2 shows the keyword co-occurrence network constructed from standardized author-provided keywords. As expected, health literacy occupied the most central position in the network, with the highest degree centrality (49), betweenness centrality (.731), closeness centrality (.929), and eigenvector centrality (1.000). Because health literacy was the core search concept used to identify the literature, this dominant position should be interpreted as a feature of the search strategy rather than as an independent substantive finding. Other prominent keywords included older adults, eHealth literacy, self-efficacy, health behavior, and self-care.
Community detection using the Louvain algorithm identified six keyword clusters (modularity=.191). Cluster 1 represented aging, chronic disease, and health management; Cluster 2, eHealth literacy, health information, and health promotion; Cluster 3, general and population-based health literacy; Cluster 4, digital and mental health literacy with psychosocial factors; Cluster 5, oral health literacy; and Cluster 6, clinical competence and patient-centered care. Representative keywords and the thematic interpretation of each cluster are shown in
Supplementary Table 1.
The supplementary network that excluded health literacy provided a clearer view of co-occurrence relationships among the remaining keywords, with older adults and eHealth literacy appearing visually prominent (
Supplementary Figure 2). After health literacy was removed, the major thematic groupings remained visually identifiable, allowing relationships among secondary research concepts to be examined without the dominance of the core search term.
Table 1 further characterizes the structural importance of individual keywords by presenting centrality measures for the top 20 keywords. Consistent with the search strategy, health literacy had the highest values across all centrality measures. Among the remaining keywords, older adults, eHealth literacy, self-efficacy, health behavior, and self-care had the highest degree centrality, indicating their prominence within the secondary keyword structure.
3. Topic Modeling
1) Topic model selection
LDA models with 3–10 topics were compared using model-fit indices, semantic coherence, topic distinctiveness, cross-seed stability, and substantive interpretability (
Supplementary Table 2). Across 10 random seeds, mean perplexity decreased from 516 at k=3 to 399 at k=10, whereas semantic coherence and cross-seed stability did not improve monotonically as k increased. Mean cross-seed stability was .488 at k=3, .445 at k=5, and .464 at k=6, and it declined for most higher-dimensional solutions. Inspection of adjacent solutions indicated that k=3 merged substantively distinct domains, whereas k=6 and higher increasingly subdivided related themes. Taken together, these criteria supported selection of the five-topic model as a parsimonious and interpretable thematic solution. The abstract-only sensitivity corpus (
N=633) showed a broadly comparable model-fit trajectory (
Supplementary Figure 3).
2) Identification and interpretation of topics
The five-topic model identified five major thematic domains in Korean health literacy research. Topic interpretation was based on the highest-probability terms and review of the 10 documents with the highest posterior probabilities for each topic.
Topic 1, Health Literacy Measurement, represented studies assessing and measuring health literacy across diverse populations. High-probability terms included “education,” “adults,” “age,” “women,” “students,” “lower,” “school,” “income,” “children,” and “functional health literacy.” Review of the 10 documents with the highest posterior probabilities showed a strong concentration of instrument development, reliability and validity testing, cross-cultural adaptation, and population-specific health literacy assessment, supporting the topic label.
Topic 2, Mental Health Literacy, centered on mental health literacy and related psychosocial and health-status factors. High-probability terms included “older adults,” “mental health,” “status,” “social support,” “mental health literacy,” “community,” “quality of life,” “middle,” “chronic disease,” and “cancer.” Representative documents addressed depression, suicide literacy, stigma, help-seeking, social support, and psychological well-being in various population groups.
Topic 3, Health Information and Communication, encompassed research on access to, use of, and communication of health and medical information. High-probability terms included “information,” “health information,” “medical,” “healthcare,” “social,” “digital,” “ability,” “service,” “public,” and “communication.” Representative documents addressed health-information seeking, information channels and services, personal health information, communication, and related policy or conceptual issues.
Topic 4, Digital Health Literacy, represented studies linking eHealth and digital health literacy with health behaviors and health-promotion contexts. High-probability terms included “eHealth literacy,” “health behavior,” “efficacy,” “health promotion behavior,” “physical activity,” “digital health literacy,” “positive,” “perceived,” “behaviors,” and “nursing students.” Representative documents frequently involved nursing or university students and examined digital information use, health-promoting behavior, and related competencies.
Topic 5, Disease-Specific Health Literacy, focused on applications of health literacy to patient care, disease-specific knowledge, education, and self-management. High-probability terms included “care,” “patients,” “knowledge,” “oral health,” “oral health literacy,” “diabetes mellitus,” “education,” “management,” “intervention,” and “educational.” Representative documents addressed oral health, diabetes, hypertension, stroke, heart failure, and other disease-specific care contexts, supporting interpretation of this topic as broader disease-specific health literacy rather than oral health literacy alone.
Figure 3 shows the 10 terms with the highest topic-word probabilities (β) for each topic, using reader-friendly labels rather than tokenized forms. The refined preprocessing step removed non-substantive methodological and reporting terms. Overall, the five topics represented health literacy measurement, mental health literacy, health information and communication, digital health literacy, and disease-specific health literacy.
3) Topic prevalence
Figure 4 shows the dominant-topic distribution for the effective full LDA corpus (
N=636). Disease-Specific Health Literacy (topic 5) was the largest topic, accounting for 21.4% of documents (n=136), followed closely by Digital Health Literacy (topic 4; 21.2%, n=135) and Mental Health Literacy (topic 2; 21.1%, n=134). Health Information and Communication (topic 3) accounted for 19.2% of documents (n=122), and Health Literacy Measurement (topic 1) accounted for 17.1% (n=109). The five topics were relatively evenly distributed, with no single domain dominating the corpus.
DISCUSSION
This study used keyword analysis, network analysis, and topic modeling to examine the knowledge structure and thematic characteristics of health literacy research in Korea. Together, the findings characterize the major themes, keyword relationships, and latent topic structure of this research field.
First, the diversity of frequently occurring keywords indicates that health literacy research in Korea extends beyond basic health information comprehension to include older populations, digital health contexts, psychosocial determinants, and health-related outcomes. This pattern is consistent with contemporary conceptualizations of health literacy as a multidimensional construct that integrates the cognitive, social, and behavioral competencies needed for effective health management [
1,
4]. The prominence of self-efficacy and social support further suggests that Korean health literacy research has emphasized patient empowerment and social determinants of health, both of which are important for health outcomes and self-management [
5,
30]. Overall, these findings indicate that Korean health literacy research treats health literacy not merely as the ability to understand health information but as a resource shaped by individual capabilities and social contexts.
Second, keyword network analysis showed that health literacy occupied a central position in the research landscape and was connected with health behavior, health promotion, chronic disease management, and psychosocial constructs. Its close associations with chronic disease-related keywords, such as diabetes mellitus and hypertension, are consistent with evidence that health literacy is relevant to self-care, treatment adherence, and disease management [
7,
9]. Because health literacy was the core search concept, its hub-like position should be interpreted cautiously. Even so, the network demonstrated meaningful links between health literacy and behavioral, clinical, and psychosocial constructs.
The presence of quality of life and social support among the top 20 keywords in the centrality analysis further underscores their relevance within the Korean health literacy research landscape and is consistent with evidence linking health literacy to quality of life [
11,
31]. These findings also fit socioecological perspectives that conceptualize health literacy as operating within interpersonal and community contexts rather than solely as an individual characteristic [
1,
4]. Similar patterns were observed in the supplementary network excluding health literacy, where older adults, self-efficacy, eHealth literacy, and social support remained visually prominent (
Supplementary Figure 2). This supplementary analysis allowed secondary keyword relationships to be examined without the visual dominance of the core search term.
Third, topic modeling identified five complementary domains: Health Literacy Measurement, Mental Health Literacy, Health Information and Communication, Digital Health Literacy and Disease-Specific Health Literacy. The relatively even topic prevalence (17.1%–21.4%) suggests that Korean health literacy research is distributed across multiple methodological and substantive domains rather than concentrated in a single area. Review of the representative documents further supported the conceptual separation of these five domains.
The Health Literacy Measurement topic highlights the methodological foundation of the literature. The documents with the highest posterior probabilities were concentrated in instrument development, psychometric evaluation, cross-cultural adaptation, and population-specific assessment. This finding indicates sustained attention in Korean health literacy research to how health literacy is operationalized and measured across diverse populations, which is a prerequisite for valid comparison and intervention evaluation.
Mental Health Literacy and Digital Health Literacy emerged as distinct substantive domains. Documents representing the mental health topic addressed depression, suicide literacy, stigma, help-seeking, social support, and psychological well-being, underscoring the relevance of mental health literacy to community-based health literacy research [
32]. The digital health topic linked eHealth and digital health literacy with health behaviors, health promotion, and student populations. This pattern is consistent with evidence that digital health literacy contributes to effective engagement with technology-mediated health information and services [
12,
19,
33-
35].
Health Information and Communication formed a separate domain centered on information access, information seeking, services, and health care communication, reflecting the information-navigation dimension of health literacy. Disease-Specific Health Literacy, the largest topic by a small margin, encompassed patient care, disease-related knowledge, education, oral health, diabetes, and self-management. The inclusion of oral health terms within this broader disease-specific topic helps explain the relatively peripheral position of individual oral-health keywords in the co-occurrence network. The keyword network captured pairwise relationships among author-provided keywords, whereas LDA identified document-level latent semantic patterns from the combined abstract-and-keyword corpus. Thus, oral health is best interpreted as one component of a broader disease-specific health literacy domain rather than as an independent dominant topic [
7,
9].
Taken together, these findings indicate that Korean health literacy research includes methodological development as well as psychosocial, informational, digital, and disease-specific applications. The relatively even thematic distribution suggests the need for context-specific strategies across multiple domains rather than a focus on a single area. For nursing practice and education, the findings point to the importance of valid health literacy assessment, effective health communication, digital health literacy, mental health literacy, and disease-specific self-management support across diverse populations.
A major strength of this study is its integration of keyword-based analyses with topic modeling based on combined abstracts and standardized author keywords. The large corpus of KCI-indexed studies further strengthens the comprehensiveness of the findings within the KCI-indexed literature. Nevertheless, several limitations should be acknowledged. First, the study relied exclusively on KCI-indexed publications; therefore, relevant studies indexed in international databases or published in nonindexed domestic sources may not have been captured. Second, the analyses used abstracts and standardized author keywords rather than full-text articles, which may have limited the depth and contextual richness of thematic extraction. Third, the findings may have been affected by preprocessing decisions, including stop-word selection, keyword standardization, thesaurus development, and document-frequency filtering. To address this concern, the preprocessing specification was finalized before sensitivity analysis and applied consistently across corpora. Fourth, topic-model selection remains inherently uncertain: log-likelihood and perplexity improved monotonically as k increased, so the five-topic solution was selected using these statistics together with semantic coherence, topic distinctiveness, cross-seed stability, and representative-document interpretability. The abstract-only sensitivity analysis yielded a broadly comparable five-topic structure, supporting the selected thematic configuration, although uncertainty remains.
Future studies that incorporate full-text analyses, additional domestic and international databases, and longitudinal designs may provide a more detailed understanding of how health literacy research in Korea has evolved. Such work may help identify emerging research priorities and support the development of evidence-based health literacy interventions and policies.
CONCLUSION
This study mapped the knowledge structure of Korean health literacy research by integrating keyword network analysis and topic modeling. The final five-topic solution identified Health Literacy Measurement, Mental Health Literacy, Health Information and Communication, Digital Health Literacy, and Disease-Specific Health Literacy as the major thematic domains, with a relatively even distribution across the corpus. These findings indicate that Korean health literacy research includes methodological development as well as psychosocial, informational, digital, and disease-specific applications. The results may help inform future research priorities and support the development of health literacy measurement strategies, interventions, nursing education programs, and evidence-based health communication initiatives tailored to diverse populations and health care contexts. Future studies should incorporate broader data sources and longitudinal perspectives to further advance health literacy research in Korea.
-
CONFLICTS OF INTEREST
The authors declared no conflict of interest.
-
AUTHORSHIP
Study conception and design - EP and KK; data curation and analysis - EP; interpretation of the data - EP and KK; and drafting or critical revision of the manuscript for important intellectual content - EP and KK.
-
FUNDING
None.
-
ACKNOWLEDGEMENT
None.
-
DATA AVAILABILITY STATEMENT
The bibliographic data analyzed in this study, including abstracts and author-provided keywords, were retrieved from academic databases and are available from the corresponding author upon reasonable request.
SUPPLEMENTARY MATERIAL
Supplementary Figure 1.
Flow diagram of the literature search and study selection process.
A total of 637 articles met the inclusion criteria and were retained for keyword-based analyses. Four articles lacked usable English abstracts; standardized author keywords were retained for the combined LDA corpus. After preprocessing and document-frequency filtering, one keyword-only record contained no retained terms, yielding an effective full LDA corpus of 636 documents. An abstract-only sensitivity corpus included 633 documents. KCI=Korean Citation Index; LDA=latent Dirichlet allocation.
kjan-2026-0510-Supplementary-Figure-1.pdf
Supplementary Figure 2.
Keyword co-occurrence network after exclusion of the keyword “health literacy.”
Node size represents keyword frequency, and colors indicate communities identified using the Louvain algorithm.
kjan-2026-0510-Supplementary-Figure-2.pdf
Supplementary Figure 3.
Perplexity across candidate latent Dirichlet allocation models (k=3–10) in the effective full corpus (N=636) and abstract-only sensitivity corpus (N=633).
The nearly overlapping trajectories indicate similar model-fit patterns across corpora; perplexity decreased progressively with increasing k and therefore was not used as the sole criterion for selecting the five-topic solution.
kjan-2026-0510-Supplementary-Figure-3.pdf
Figure 1. Word cloud of author-provided keywords after exclusion of “health literacy.“
Word size is proportional to keyword frequency.
Figure 2. Keyword co-occurrence network and Louvain community structure of Korean health literacy research.
Node size represents categorized keyword frequency (<10, 10–19, 20–49, 50–99, and ≥100 occurrences), node color represents Louvain community membership, and edge width represents keyword co-occurrence frequency. The thematic interpretation of each cluster is presented in
Supplementary Table 1.
Figure 3. Top 10 terms for each topic identified by the final five-topic latent Dirichlet allocation model.
Bar lengths represent topic-word probabilities (β). Display labels were converted from tokenized forms to reader-friendly terms. Topic labels were assigned based on the highest-probability terms and review of the 10 documents with the highest posterior topic probability for each topic.
Figure 4. Distribution of documents across the five latent topics in the effective full latent Dirichlet allocation corpus (N=636)
.Percentages represent the proportion of documents assigned to the dominant topic based on the highest posterior probability (γ). T1=Health Literacy Measurement; T2=Mental Health Literacy; T3=Health Information and Communication; T4=Digital Health Literacy; T5=Disease-Specific Health Literacy.
Table 1.Centrality Measures for the Top 20 Keywords in the Health Literacy Network
|
Rank |
Keyword |
Degree |
Betweenness |
Closeness |
Eigenvector |
Frequency |
|
1 |
Health literacy |
49 |
.731 |
.929 |
1.000 |
436 |
|
2 |
Older adults |
23 |
.076 |
.627 |
.659 |
100 |
|
3 |
eHealth literacy |
17 |
.095 |
.598 |
.514 |
63 |
|
4 |
Self-efficacy |
12 |
.009 |
.553 |
.491 |
41 |
|
5 |
Health behavior |
11 |
.014 |
.547 |
.401 |
51 |
|
6 |
Self-care |
10 |
.006 |
.531 |
.383 |
32 |
|
7 |
Social support |
9 |
.007 |
.536 |
.375 |
35 |
|
8 |
Nursing students |
9 |
.006 |
.536 |
.338 |
28 |
|
9 |
Health promotion |
8 |
.004 |
.531 |
.342 |
31 |
|
10 |
Diabetes mellitus |
8 |
.002 |
.520 |
.334 |
31 |
|
11 |
Quality of life |
7 |
.001 |
.515 |
.326 |
30 |
|
12 |
Literacy |
6 |
.003 |
.510 |
.162 |
25 |
|
13 |
Knowledge |
6 |
.001 |
.510 |
.284 |
24 |
|
14 |
Digital health literacy |
6 |
.002 |
.510 |
.267 |
23 |
|
15 |
Health promotion behavior |
6 |
.001 |
.520 |
.272 |
21 |
|
16 |
Mental health literacy |
5 |
.002 |
.505 |
.186 |
35 |
|
17 |
Health information |
5 |
.000 |
.515 |
.225 |
18 |
|
18 |
Oral health |
5 |
.002 |
.505 |
.169 |
16 |
|
19 |
Hypertension |
5 |
.000 |
.505 |
.239 |
15 |
|
20 |
Digital health |
4 |
.000 |
.510 |
.230 |
20 |
REFERENCES
- 1. Nutbeam D, Lloyd JE. Understanding and responding to health literacy as a social determinant of health. Annu Rev Public Health. 2021;42:159-73. https://doi.org/10.1146/annurev-publhealth-090419-102529
- 2. World Health Organization. Health literacy development for the prevention and control of noncommunicable diseases: Volume 1. Overview [Internet]. Geneva: World Health Organization; 2022 [cited 2026 April 16]. Available from: https://www.who.int/publications/i/item/9789240055339
- 3. Centers for Disease Control and Prevention. What is health literacy? [Internet]. Atlanta, GA: Centers for Disease Control and Prevention; 2024 [cited 2026 April 16]. Available from: https://www.cdc.gov/health-literacy/php/about/index.html
- 4. Sorensen K, Levin-Zamir D, Duong TV, Okan O, Brasil VV, Nutbeam D. Building health literacy system capacity: a framework for health literate systems. Health Promot Int. 2021;36(Suppl 1):i13-23. https://doi.org/10.1093/heapro/daab153
- 5. Magi CE, El Aoufy K, Amato C, Longobucco Y, Bambi S, Vellone E, et al. The association between self-care and health literacy in patients with chronic diseases: a systematic review and meta-analysis. J Clin Nurs. 2026;35(7):3011-29. https://doi.org/10.1111/jocn.70291
- 6. Sheehan S, Bernues-Caudillo L, de Brun A, Drury A. Health literacy, self-management and patient-reported outcomes in prostate cancer survivors: a mixed methods systematic review. Semin Oncol Nurs. 2026;42(1):152056. https://doi.org/10.1016/j.soncn.2025.152056
- 7. Yu J, Shi N, Zhao J, Li F, Wang J, Jin M, et al. Effectiveness of health literacy interventions for blood pressure control, medication adherence, and self-efficacy in older adults with hypertension: a systematic review and meta-analysis. J Eval Clin Pract. 2026;32(2):e70419. https://doi.org/10.1111/jep.70419
- 8. Marshall N, Butler M, Lambert V, Timon CM, Joyce D, Warters A. Health literacy interventions and health literacy-related outcomes for older adults: a systematic review. BMC Health Serv Res. 2025;25(1):319. https://doi.org/10.1186/s12913-025-12457-7
- 9. Magnani JW, Mujahid MS, Aronow HD, Cene CW, Dickson VV, Havranek E, et al. Health literacy and cardiovascular disease: fundamental relevance to primary and secondary prevention: a scientific statement from the American Heart Association. Circulation. 2018;138(2):e48-74. https://doi.org/10.1161/CIR.0000000000000579
- 10. Bae EJ, Yoon JY. Health literacy as a major contributor to health-promoting behaviors among Korean teachers. Int J Environ Res Public Health. 2021;18(6):3304. https://doi.org/10.3390/ijerph18063304
- 11. Zheng M, Jin H, Shi N, Duan C, Wang D, Yu X, et al. The relationship between health literacy and quality of life: a systematic review and meta-analysis. Health Qual Life Outcomes. 2018;16(1):201. https://doi.org/10.1186/s12955-018-1031-7
- 12. Xie G, Liao J, Tang X, Yang Y, Han F, Liu D, et al. Digital health literacy in medical education: a scoping review of current challenges and development strategies. BMC Med Educ. 2026;26(1):567. https://doi.org/10.1186/s12909-026-08903-7
- 13. Pelizzari N, Covolo L, Ceretti E, Fiammenghi C, Gelatti U. Defining, assessing, and implementing organizational health literacy: barriers, facilitators, and tools: a systematic review. BMC Health Serv Res. 2025;25(1):599. https://doi.org/10.1186/s12913-025-12775-w
- 14. Faerevaag FS, Kalsnes B, Tennfjord MK, Molin M. Mapping the landscape of critical health literacy: a comprehensive scoping review of research trends and associations with health behaviors. J Health Commun. 2026;31(1):34-54. https://doi.org/10.1080/10810730.2025.2608161
- 15. Hovingh JW, Elderson-van Duin C, Kuipers DA, van Rood Y, Ludden GDS, Hanssen DJC, et al. Tailoring for health literacy in the design and development of eHealth interventions: systematic review. JMIR Hum Factors. 2025;12:e76172. https://doi.org/10.2196/76172
- 16. Bakhtiarvand SZ, Rahaei Z, Sadeghian HA, Fatehi F, Soltani S, Zareiyan A, et al. The constructs of health literacy in children: a systematic review. BMC Public Health. 2025;25(1):3352. https://doi.org/10.1186/s12889-025-24573-4
- 17. Kang SJ, Lee MS. Evidence-based health literacy improvements: trends on health literacy studies in Korea. Korean J Health Educ Promot. 2015;32(4):93-108. https://doi.org/10.14367/kjhep.2015.32.4.93
- 18. Lee M, Shin HG, Lee MJ, Park CY. Research trends and policy issues of health literacy in Korea. J Health Tech Assess. 2018;6(1):22-32. https://doi.org/10.34161/johta.2018.6.1.004
- 19. Kim K, Shin S, Kim S, Lee E. The relation between eHealth literacy and health-related behaviors: systematic review and meta-analysis. J Med Internet Res. 2023;25:e40778. https://doi.org/10.2196/40778
- 20. Yoon J, Lee M, Cho J. Concept of digital health literacy: a scoping review. Public Health Weekly Report. 2024;17(48):2095-133. https://doi.org/10.56786/PHWR.2024.17.48.2
- 21. Park S. Analysis of domestic health literacy research trends using topic modeling: focusing on papers published in journals from 2012 to 2021. J Humanit Soc Sci 21. 2023;14(3):1185-200. https://doi.org/10.22143/HSS21.14.3.83
- 22. Wang J, Shahzad F. A visualized and scientometric analysis of health literacy research. Front Public Health. 2022;9:811707. https://doi.org/10.3389/fpubh.2021.811707
- 23. Paucar-Caceres A, Vilchez-Roman C, Quispe-Prieto S. Health literacy concepts, themes, and research trends globally and in Latin America and the Caribbean: a bibliometric review. Int J Environ Res Public Health. 2023;20(22):7084. https://doi.org/10.3390/ijerph20227084
- 24. Park J, Han AY. Medication safety education in nursing research: text network analysis and topic modeling. Nurse Educ Today. 2023;121:105674. https://doi.org/10.1016/j.nedt.2022.105674
- 25. Won J, Kim K, Sohng KY, Chang SO, Chaung SK, Choi MJ, et al. Trends in nursing research on infections: semantic network analysis and topic modeling. Int J Environ Res Public Health. 2021;18(13):6915. https://doi.org/10.3390/ijerph18136915
- 26. Blondel VD, Guillaume JL, Lambiotte R, Lefebvre E. Fast unfolding of communities in large networks. J Stat Mech Theory Exp. 2008;2008(10):P10008. https://doi.org/10.1088/1742-5468/2008/10/P10008
- 27. Csardi G, Nepusz T. The igraph software package for complex network research. Inter Journal Complex Syst. 2006;1695:1-9.
- 28. Blei DM, Ng AY, Jordan MI. Latent Dirichlet allocation. J Mach Learn Res. 2003;3:993-1022.
- 29. Griffiths TL, Steyvers M. Finding scientific topics. Proc Natl Acad Sci U S A. 2004;101(Suppl 1):5228-35. https://doi.org/10.1073/pnas.0307752101
- 30. Nam HJ, Yoon JY. Pathways linking health literacy to self-care in diabetic patients with physical disabilities: a moderated mediation model. PLoS One. 2024;19(3):e0299971. https://doi.org/10.1371/journal.pone.0299971
- 31. Ehmann AT, Groene O, Rieger MA, Siegel A. The relationship between health literacy, quality of life, and subjective health: results of a cross-sectional study in a rural region in Germany. Int J Environ Res Public Health. 2020;17(5):1683. https://doi.org/10.3390/ijerph17051683
- 32. Oguntoye OV, Afolayan JA, Oguntoye M, Ajala DE, Esan DT. Mental health literacy programs in Nigeria: a comprehensive scoping review of interventions, challenges and outcomes. Arch Psychiatr Nurs. 2026;61:152078. https://doi.org/10.1016/j.apnu.2026.152078
- 33. Kim S, Park C, Park S, Kim DJ, Bae YS, Kang JH, et al. Measuring digital health literacy in older adults: development and validation study. J Med Internet Res. 2025;27:e65492. https://doi.org/10.2196/65492
- 34. Kang H, Baek J, Chu SH, Choi J. Digital literacy among Korean older adults: a scoping review of quantitative studies. Digit Health. 2023;9:20552076231197334. https://doi.org/10.1177/20552076231197334
- 35. Arias Lopez MDP, Ong BA, Borrat Frigola X, Fernandez AL, Hicklent RS, Obeles AJT, et al. Digital literacy as a new determinant of health: a scoping review. PLOS Digit Health. 2023;2(10):e0000279. https://doi.org/10.1371/journal.pdig.0000279