Plastic SurgeryLarge Language ModelsMedical TourismSouth KoreaSentiment Analysis

Understanding International Patient Experiences in South Korean Plastic Surgery: An AI-Assisted Analysis of Patient Reviews, Service Quality, Communication, and Trust

Ishaan Rastogi Published August 15, 2026 CC-BY

South Korea is one of the most visible destinations in global cosmetic-surgery tourism, and the experience patients report there depends on far more than the surgical result: communication, consultation quality, pricing, staff interaction, and post-operative support all shape how a procedure is remembered and described. This paper reports a mixed-methods study of international patient experience in South Korean plastic surgery, built around a dataset of 74 patient narratives and service ratings across 23 clinics, assembled in three stages: a hand-compiled research document (9 clinics, 36 observations), a Google Maps-sourced expansion (6 clinics, 20 observations), and a deliberately non-Google expansion (8 clinics, 18 observations) drawn from a medical-tourism marketplace (WhatClinic), a Korean-language patient community (Gangnam Unni), a Korean booking/discount platform (Babitalk), a Korean review aggregator (Sungyesa), a medical-tourism referral marketplace (Bookimed), and a public social-media post (Threads). Every clinic in the combined dataset has at least three of five service dimensions populated by real review content. Price remained the most consistently low-scoring dimension (lowest or tied-lowest at 11 of 22 comparable clinics), while staff interaction and post-operative care were the most consistently praised. The three-stage collection design produced a direct, measured finding about data-collection method itself: the Google-sourced batch contained zero mixed-sentiment observations, while the non-Google batch recovered three, suggesting the earlier "no mixed sentiment" pattern was a property of Google's specific review-panel selection rather than of web-sourced data in general. The expansion also surfaced a substantive negative review alleging undisclosed additional charges before anesthesia, a documented historical "ghost doctor" controversy independently corroborated against a 2016 Korea Times report[2], and, via Bookimed[3], an unusually long first-person account of an entire multi-week patient journey that included a frightening but ultimately normal post-operative complication. An AI-assisted lexicon classifier was run against all 74 narratives and checked against human-coded labels, reaching 82% sentiment accuracy and a macro F1 of 0.74 on theme classification; a specific, honestly-reported failure mode emerged where the classifier's keyword-based "review credibility" trigger over-fires on the phrase "verified reviewer," a phrasing artifact of how the newer batches were written up rather than a deeper semantic problem. The paper does not rank clinics or make claims about surgical safety; its contribution is a research-based account of what shapes international patient experience, a demonstration of how data-source diversification changes what a review dataset can say, and an assessment of where AI-assisted analysis helps and where it needs human oversight.

Authors: Ishaan Rastogi1^1, Prakriti Puri2^2

1^1 Department of Computer Science Engineering, Amity School of Engineering & Technology, Amity University Uttar Pradesh

2^2 Department of Computer Science Engineering, Amity School of Engineering & Technology, Amity University Uttar Pradesh

1. Introduction

Every year, hundreds of thousands of international patients travel to South Korea for cosmetic and reconstructive procedures, drawn by the country's concentration of clinics, its reputation for aesthetic outcomes, and its position at the center of the "Korean Look" associated with exported Korean popular culture[4]. Government figures cited in reviews of the sector describe a market that grew for years before the pandemic, with cosmetic surgery consistently the highest-earning category of foreign-patient treatment and Gangnam functioning as its geographic and commercial hub[5]. Most existing research on this market, however, approaches it from the clinic's side: what drives a hospital's ability to attract and retain international patients, and which service-quality dimensions predict satisfaction and loyalty[6,7,8]. Comparatively little work has started from the other direction, using what patients themselves write in reviews and narratives as the primary evidence, and almost none of it treats which platform a review comes from as a variable worth measuring rather than an incidental detail of data collection.

That second gap matters as much as the first. A translator who was helpful in one appointment and unavailable in the next, a price that was quoted one way at consultation and changed before surgery, a recovery that felt frightening at two weeks and fine at six months. Narratives like these are rich but slow to work with one at a time, and they differ depending on where a researcher looks for them. A Google Maps listing's default panel, a medical-tourism marketplace's verified-review system, and a Korean-language patient community each select and surface reviews differently, and a dataset built from only one of them risks mistaking a property of the platform for a property of the clinics on it.

This paper, prepared under the K-PSI research proposal[9], began with a preliminary dataset compiled by the research team, entirely from publicly available internet sources, covering patient reviews and clinic-level ratings across nine South Korean plastic surgery clinics. This revised version extends that dataset through two further passes of the same kind of internet research. The second pass added six clinics from public Google Maps listings. Once that pass's own results showed a measurable positivity skew relative to the first pass, the third pass deliberately sourced eight more clinics from six different non-Google platforms, specifically to test whether that skew was a property of Google's panel or of internet-sourced review data generally. The paper combines four things: a thematic reading of the patient narratives, a comparison of the numerical service ratings across five dimensions (language and communication, staff, booking and consultation, price, and post-operative care), an AI-assisted analysis of sentiment and themes evaluated against human coding, and an explicit, three-way comparison across the three collection passes. Consistent with the proposal's own framing, the goal is not to rank clinics or to evaluate medical quality. It is to understand what characterizes the reported experience of international patients in this setting, and to assess honestly what AI-assisted analysis adds to that understanding, what data-source diversity adds to it, and where each falls short.

Figure 1.1. The kind of procedure vocabulary (rhinoplasty, facial contouring, chin surgery) that recurs throughout the patient narratives analyzed in this paper, shown here for orientation only via a generic marketing-style graphic. It is not drawn from, and does not represent, any patient or clinic in the dataset.

2. Literature Review

2.1 Service quality and satisfaction in Korean cosmetic medical tourism

The most directly relevant prior work comes from studies of Chinese patients, who make up a large share of the international cosmetic-surgery market in Korea. Park et al.[6] surveyed 334 Chinese patients who had undergone cosmetic surgery in Korea and used structural equation modeling to show that a hospital's service quality along the dimensions of tangibles, assurance, and empathy shaped patients' attitudes toward medical tourism, which in turn predicted satisfaction; the same study found that the service quality of facilitating agents, the intermediaries who arrange travel, translation, and appointments, moderated how strongly hospital service quality translated into satisfaction. This is one of the few studies to treat the agent relationship as a distinct variable rather than folding it into the hospital experience, and it lines up closely with what appears in the reviews analyzed here, where translators and coordinators are named almost as often as clinics themselves, across every platform this dataset draws from.

Anvarjonov and Um[8] extended this line of work to the digital side of the relationship, finding that a clinic's social-media marketing capability improved international patient satisfaction largely by reducing perceived risk, that is, by making it easier for a prospective patient to find credible information before committing to travel. Related survey work on Chinese patronage patterns found that hospital size and social-media activity, rather than location in Gangnam specifically, predicted the volume of Chinese patients a clinic attracted[10], and a loyalty study of 158 Chinese patients found that medical and tourism service quality together explained roughly 85% of the variance in customer loyalty[7]. Two qualitative and interpretive studies add texture that quantitative surveys miss: Holliday et al.[4] describe how Korean surgeons and Chinese patients sometimes hold different understandings of what the surgery is meant to achieve, situating this in a broader discourse of "medical nationalism," and Koh's[5] review documents the institutional infrastructure, medical visas, accreditation, tax refunds, that the Korean government built around the sector as it grew through the 2010s.

Service-quality research in medical tourism more broadly has relied heavily on the SERVQUAL framework, which measures the gap between what patients expect and what they perceive they received across five dimensions: reliability, assurance, tangibles, empathy, and responsiveness. A systematic review and meta-analysis of SERVQUAL studies in Asian healthcare settings, including one study conducted in South Korea, found negative expectation-perception gaps across nearly every dimension studied, with the size of the gap varying most by reliability and tangibles[11]. Applications of SERVQUAL specifically to medical tourists[12,13] and, more recently, structural models linking service quality to trust and revisit intention in other medical-tourism destinations[14] consistently find that trust, not satisfaction alone, is what predicts whether a patient would return or recommend the destination to others. None of this work, however, analyzes patient narratives directly; it relies on structured survey instruments administered after the fact, which is precisely the gap a review-based, narrative-driven study can address.

2.2 Online reviews and NLP-based sentiment analysis in surgical care

A separate but adjacent literature has used natural language processing to analyze online reviews of surgeons and surgical experiences, mostly outside the cosmetic-tourism context. Several studies applying sentiment analysis to physician-review sites, spine surgery[15], orthopaedic practice[16], vascular surgery[17], pediatric orthopaedics[18], hand surgery[19], and refractive surgery[20], have converged on a similar finding: patients writing online reviews of surgeons focus disproportionately on non-clinical factors, communication, bedside manner, staff friendliness, and pain management, rather than technical outcome, and these non-clinical themes are strong predictors of whether a review reads as positive or negative. Within cosmetic surgery specifically, a study of RealSelf reviews of blepharoplasty patients categorized narratives into recurring factors such as aesthetic outcome, comfort, cost, and provider manner[21], and Ramnarine[22] applied unsupervised topic modeling and sentiment analysis to plastic-surgery social-media posts as a methodological demonstration of how these techniques scale to large, unlabeled text corpora. More recent work using deep-learning sentiment models on RealSelf data for gender-affirming top surgery found that sentiment scores correlated strongly with patients' self-reported "worth it" ratings, and that procedure cost itself was not a significant predictor of satisfaction once outcome-related sentiment was accounted for[23]. Taken together, this body of work establishes that sentiment analysis of surgical reviews is methodologically workable and tends to surface the same handful of experiential themes, communication, staff behavior, cost, and outcome, across very different surgical specialties and, this paper's expanded dataset suggests, across very different review platforms as well.

2.3 Review bias, incentivized reviews, and trust

Because the K-PSI dataset includes reviews in which patients explicitly disclose receiving a discount in exchange for writing about their experience, the literature on incentivized and biased online reviews is directly relevant. Two related strands of work reach a similar conclusion from different directions. One strand shows that unsolicited, voluntary reviews tend to be more extreme than the underlying distribution of customer experience, because people with strong opinions, positive or negative, are more motivated to write unprompted, and that incentives or solicitations can partially correct this by drawing out reviews from people with moderate, unremarkable experiences[24,25]. The other strand focuses on how incentives affect the reader's trust rather than the review's statistical distribution: incentivized reviews are read as less credible than organic ones because a monetary or in-kind reward creates a perceived conflict of interest between the reviewer and the business[26,27].

Within healthcare specifically, the evidence on review bias is less about incentives and more about platform mechanics. Kordzadeh[28] compared physician ratings posted on hospitals' own websites against ratings for the same physicians on independent commercial platforms and found that hospital-hosted ratings were systematically higher and less dispersed. That is a between-platform bias that this paper's expanded dataset now demonstrates directly rather than only citing: Section 5.3 shows the same clinics' review pool reading measurably differently depending on whether it was drawn from Google Maps or from a mix of marketplace, community, and referral platforms. Set against this, negativity-bias research suggests that readers give disproportionate diagnostic weight to negative review content regardless of how experienced or careful they are as consumers[29], directly relevant to how the Grand Plastic Surgery historical-controversy review (Section 5.4) should be read.

2.4 AI-assisted coding and its validation against human judgment

The proposal's methodology calls for validating AI-assisted classification against human-coded data using precision, recall, and F1-score, which places this study within a growing body of work assessing how well large language models perform at qualitative coding tasks traditionally done by trained human researchers. A comparison of LLM output against expert-consensus coding of healthcare focus-group transcripts found that LLM-assisted deductive coding achieved agreement statistically indistinguishable from blinded human coders[30]. GPT-4 compared against three trained human reviewers on sentiment polarity and thematic classification of patient interview transcripts achieved a Cohen's kappa of 0.69[31]. A larger comparison of seven LLMs against 33 human annotators found models generally matched or exceeded human inter-rater reliability on sentiment, but that both humans and models performed poorly at detecting sarcasm[32]. This paper's own classifier failure mode (Section 5.5), a keyword trigger over-firing on an incidental phrasing convention, is a small, concrete instance of a broader point this literature makes: classifier performance is highly sensitive to surface features of exactly the text register it is run against, not a fixed property of the method.

2.5 Positioning of the present study

Read together, this literature leaves a specific gap that the present study is positioned to address, and the two-round expansion in this revision adds a second, more practical one: almost none of the literature above treats data-collection method as an experimental variable. Research on Korean cosmetic medical tourism has established which service-quality dimensions predict satisfaction and loyalty, but almost entirely through structured surveys. Research on sentiment analysis of surgical reviews has shown that NLP methods reliably extract communication, staff, and cost-related themes from patient text, but has largely been applied to single-platform corpora. K-PSI's expanded version sits at this intersection: it uses a small, structurally rich, three-batch dataset (one hand-compiled, one Google-sourced, one deliberately multi-platform) to ask not only what these bodies of work would predict about patient experience, but what changes, measurably, when the same research question is asked of six different kinds of platform.

3. Research Problem and Research Questions

Existing research establishes that service quality, information delivery, perceived risk, and satisfaction are connected in cosmetic medical tourism, but it has consistently accessed this relationship through post hoc surveys rather than through the narratives patients generate on their own terms, and almost never through more than one platform at a time. Four research questions guided the analysis, unchanged from the original proposal:

RQ1. What factors most frequently characterize positive and negative patient experiences associated with plastic-surgery services in South Korea?

RQ2. How are communication, staff interaction, consultation and booking, pricing, and post-operative care related to reported overall satisfaction?

RQ3. What recurring themes related to trust, information accessibility, language barriers, pricing, and review credibility appear in patient narratives?

RQ4. How accurately can an AI-assisted NLP approach identify sentiment and key themes in the collected patient narratives compared with human coding?

RQ2 is addressed only indirectly: an explicit overall-satisfaction figure was recorded for just one clinic in the original source material, and deriving one as the mean of the other five dimensions for the remaining clinics turned out to be circular rather than independent information (documented and removed in an earlier iteration of this codebase). RQ2 is addressed here only through the descriptive relationships between the five dimensions themselves (Sections 5.1 and 5.2).

4. Methodology

4.1 Research design

This study uses a mixed-methods exploratory design combining five components: a qualitative thematic analysis of patient narratives, a quantitative comparison of numerical service ratings, a computational component in which an NLP/LLM pipeline assists with classification and is evaluated against human coding, a three-way comparison across the hand-compiled, Google-sourced, and non-Google-sourced batches, and a minimum-data-coverage rule applied uniformly across every clinic regardless of source. This design follows directly from the K-PSI research proposal and remains appropriate given the exploratory, preliminary nature of the dataset.

4.2 Data

Every observation in this dataset, across all three collection passes, was produced by desk research on publicly available internet sources: reading review listings, marketplace profiles, and community posts, and paraphrasing what was found there. At no point did the research team contact a patient or a clinic directly, conduct an interview, or administer a survey; there is no primary fieldwork in this study, only secondary compilation and analysis of what patients had already, independently, chosen to post publicly. The three passes described below differ in which internet sources they drew on and when they were conducted, not in kind.

The first pass produced the original nine-clinic document[33], referred to by the pseudonyms used in that material (DA, Girin, View, Oneul, AB, Chiu, Braun, I-one, and HighEnd). It was compiled by the research team from a mix of internet sources per clinic (Google Reviews, RealSelf, Babitalk, UNNI, and Sungyesa, together with Korean-language community discussion), condensed into four patient observations per clinic across five service dimensions. An explicit overall-satisfaction rating was recorded for only one clinic (DA), and the clinic-level average figures reported for Oneul in this material are byte-for-byte identical to View's, an almost-certain data-entry duplication; Oneul is excluded from the cross-clinic comparisons reported below.

The second pass added six clinics (IDHospital, JW, Banobagi, Wonjin, Grand, and Dream) by manually browsing each clinic's public Google Maps[34] listing on 2026-08-15 and reading whatever reviews the listing's default panel surfaced. This was non-automated browsing: no bulk harvesting, pagination-looping, or login-bypassing. Every observation was paraphrased and anonymized, rating only the dimensions the review text actually discussed. This pass's own results (Section 5.3) showed it read measurably more positive and less mixed-sentiment than the first pass, which motivated a third.

The third pass added eight more clinics specifically by searching outside Google, across six further internet sources:

  • WhatClinic[35] (a medical-tourism marketplace with independently verified patient reviews): Pitangui Medical & Beauty, Hyundai Mihak Plastic Surgery, and Bbae Clinic, plus two additional reviews backfilled onto Grand Plastic Surgery (from the second pass) to bring its dimension coverage up to the same minimum applied everywhere else.
  • Gangnam Unni[36] (강남언니, a Korean-language plastic-surgery patient community and booking platform): Woo-a Plastic Surgery, Kai Plastic Surgery, and Teia Clinic, each from one or two community posts read and paraphrased directly from Korean.
  • Babitalk[37] (바비톡, a Korean discount/booking platform for aesthetic procedures) and Sungyesa[38] (성예사, the same Korean review-aggregation platform already used for Chiu in the first pass), together with one Threads[39] post: all four sources converged on a single clinic, Lala Plastic Surgery, making it the most multi-sourced clinic in the dataset.
  • Bookimed[3] (a medical-tourism referral marketplace): VG Plastic Surgery, via two verified reviews for surgeon Im Young Min, including one unusually long, detailed first-person account of an entire multi-week patient journey.

Every observation in this pass, like the second, was paraphrased into the existing narrative-summary schema, anonymized (no reviewer names, photos, or profile detail), and rated only on dimensions the source text actually discussed. Korean-language sources (Gangnam Unni, Babitalk, Sungyesa) were read and translated directly by the authors; translation judgment calls are the authors' own and are not independently verified by a second reader, which is disclosed as a limitation in Section 7.

A minimum-coverage rule was applied across the whole dataset, all three passes included: every clinic must have real review content covering at least three of the five service dimensions before being included. Several candidate clinics found during the third pass (on WhatClinic: Optima, Unique, VVLY; on Gangnam Unni: two other community-post subjects) were read and rejected because their available review content covered only one or two dimensions, most often because a listing had only one short review or a post focused narrowly on a single topic. Grand Plastic Surgery, retained from the second pass, initially had only one of five dimensions populated and was brought up to three via the WhatClinic backfill described above rather than being dropped.

4.3 Human thematic coding

Each of the 74 narrative observations was read and coded into one or more of the following categories: communication and language, staff interaction, booking and consultation, pricing, post-operative care, and two cross-cutting categories, trust and information accessibility, and review credibility. Sentiment was recorded as positive, negative, mixed, or neutral. The coding scheme is identical across all three passes.

4.4 AI-assisted NLP

A deterministic, keyword-driven classifier was applied to all 74 narratives to independently generate sentiment labels, thematic tags, and a review-bias flag. Its vocabulary was expanded once, generically, before validation, to cover the more casual register of web-sourced text; it was not re-tuned after seeing the third-pass results, including the specific failure mode identified in Section 5.5, so that the reported metrics reflect the classifier's actual out-of-sample behavior rather than a classifier fitted to this dataset's known quirks.

4.5 Quantitative analysis

Descriptive statistics were computed for each of the five service dimensions across the 22 comparable clinics (Oneul excluded). Dimension coverage varies by clinic because of the minimum-coverage rule (Section 4.2): some clinics have all five dimensions populated from rich, multi-review sources, while others have exactly three. Per-clinic "lowest-scoring dimension" figures are reported for every clinic, but should be read with this in mind: a clinic with only three populated dimensions is being compared on a narrower basis than one with five.

4.6 AI validation

Human-coded sentiment and theme labels were compared against the AI-generated labels for all 74 narratives using accuracy, precision, recall, and F1-score, computed by a small, purpose-built metrics implementation rather than a third-party statistical library.

5. Results

5.1 Descriptive comparison of service dimensions

Table 1 reports the mean rating (on a 5-point scale) for each of the five service dimensions, averaged across the 22 comparable clinics.

Table 1. Mean service-dimension ratings across 22 clinics

Dimension Mean Minimum (clinic) Maximum (clinic)
Post-operative care 4.45 3.00 (HyundaiMihak) 5.00 (Banobagi)
Staff 4.44 2.00 (HyundaiMihak) 5.00 (Bbae)
Booking/consultation 4.32 3.00 (Braun) 5.00 (JW)
Language/communication 4.19 2.00 (HyundaiMihak) 5.00 (AB)
Price 3.52 1.00 (Wonjin) 5.00 (Bbae)

Figure 5.1.1. Mean service-dimension ratings across 22 clinics.

Price remained the lowest-scoring dimension by a clear margin and was the single lowest-rated (or tied-lowest) dimension at 11 of the 22 comparable clinics. Its range widened substantially with the expansion: Wonjin's negative review (undisclosed additional charges before anesthesia) sets the floor at 1.00, while Bbae Clinic's WhatClinic review, explicitly praising "reasonable prices," sets a new ceiling at 5.00. The same dimension now has genuine evidence on both extremes, which the original 9-clinic sample did not.

Hyundai Mihak Plastic Surgery is the clearest illustration of why averaging across dimensions can obscure more than it reveals: its single source review is genuinely mixed (affordable price and an acceptable outcome, against a small facility, slow message response, and only 1-2 English-speaking staff), and it is simultaneously the dataset's lowest scorer on staff and language/communication. This is retained deliberately rather than smoothed into a single sentiment label, because the mix is itself the finding.

Figure 5.1.2 shows the full clinic-by-dimension matrix, including missing cells, which makes the coverage difference between richly-sourced and minimally-sourced clinics directly visible.

Figure 5.1.2. Matrix mapping out each clinic with each dimension on the basis of mean rating

5.2 Thematic findings from patient narratives

Communication support remained the most consistently praised element across all three batches, and the non-Google batch added a specific new sub-pattern: at Kai Plastic Surgery and Woo-a Plastic Surgery, two independent Gangnam Unni community posts each praised a consultation that explicitly declined to recommend a larger, more expensive option than the patient's frame warranted. That is a direct, positive counter-example to the upselling and pushy-consultation complaints that recur elsewhere in this dataset (Girin, HighEnd's contrasting praise for not aggressively recommending procedures, and now these two).

Pricing remained the most consistent source of complaint where discussed, reinforced by Wonjin's undisclosed-charges review, but the non-Google batch also produced the dataset's clearest positive price narratives: Bbae Clinic ("reasonable prices"), Teia Clinic ("reasonable for the level of care"), Dream ("VIP treatment without the VIP price"), and a Sungyesa review of Lala Plastic Surgery reporting a zero-cost consultation. Price is not uniformly negative in this dataset; it is the dimension patients discuss with the most range, from the sharpest complaints to the most specific praise.

5.3 Comparison across data-collection batches

Because every observation is tagged with the batch and platform it was drawn from, the three collection methods can be compared directly. Figure 5.3.1 shows the human-coded sentiment distribution split three ways.

Figure 5.3.1 Metrics of AI-assisted sentiment classification versus human coding.

The pattern is specific enough to support a real methodological claim, not just a general "web data is different" observation. The Google-sourced batch (6 clinics, 20 observations) contains zero mixed-sentiment observations and only 2 negative ones. The non-Google batch (10 clinics: 8 new plus the 2 backfilled Grand rows, 18 observations) contains 3 mixed-sentiment observations and, notably, zero negative ones. Diversifying away from Google did not simply make the data "more negative" or "more balanced" in a generic sense. It recovered a specific sentiment category, mixed, that had gone missing, while the small number of clearly negative accounts in this dataset (Wonjin, Grand's historical-controversy citation) both remain attributable to Google specifically. This is consistent with a more precise version of the hypothesis raised in the prior revision: it is not that "web-sourced data skews positive" in general, but that Google Maps' specific default-panel selection mechanism suppresses the kind of qualified, both-good-and-bad account that a longer-form marketplace review (WhatClinic), a first-person community post (Gangnam Unni), or a multi-week patient-journey narrative (Bookimed) is more likely to contain.

5.4 The Grand Plastic Surgery controversy

One Google-sourced observation warrants separate treatment because of its severity. A recent Google review of Grand Plastic Surgery invoked a "ghost doctor" (shadow doctor) scandal. Independent verification found this is not an isolated or unsubstantiated claim: The Korea Times reported in 2016[2] that a director of Grand Plastic Surgery in Gangnam was implicated in ghost-surgery reporting, and a 2022 Duke University Press academic article[40] references a 2014 Korea Medical Association press conference concerning the same clinic. This paper reports the review and its corroboration as a documented historical controversy from 2014 to 2016. It does not, and cannot, speak to the clinic's current-day practices; Grand's other reviews in this dataset, from Google (a tummy tuck, a blepharoplasty) and from the WhatClinic backfill (a rhinoplasty, an extensive multi-procedure case with a promptly and generously resolved post-operative concern), were all positive. The juxtaposition illustrates, concretely, the negativity-bias dynamics discussed in Section 2.3: a single serious historical allegation, even independently corroborated, coexists in the same small dataset with several ordinary recent positive experiences.

5.5 AI-assisted classification against human coding

Sentiment classification. Overall accuracy was 0.82 (macro F1 0.66) across all 74 narratives. The classifier remained strong on positive (F1 0.90, n=52) and negative (F1 0.84, n=10) narratives, and improved somewhat on mixed-sentiment narratives (F1 0.46, n=8, up from 0.29 on the 56-observation sample) now that the non-Google batch gave it more mixed examples to be right or wrong about.

Figure 5.5.1. AI-assisted them classification F1 score category-wise.

Theme classification. Micro F1 was 0.74 and macro F1 0.74. A specific, honestly-reported failure mode appeared in this round: review_credibility precision fell to 0.46 (F1 0.52, the weakest category), because nearly every WhatClinic-sourced narrative in this paper's own write-up begins with the phrase "WhatClinic-verified reviewer...", and the classifier's review_credibility lexicon includes the bare keyword "review," which matches "reviewer" as a substring. This is not a deep semantic failure; it is a phrasing artifact created by this paper's own narrative-summary convention interacting with a naive keyword matcher. Per Section 4.4, the classifier was not re-tuned after this was discovered, specifically so the reported metrics reflect its real behavior rather than a version quietly patched to look better on this dataset. It is reported here instead as a concrete illustration of exactly the surface-sensitivity the AI-validation literature describes (Section 2.4): a one-word lexicon choice, interacting with an incidental writing convention, measurably changed a reported precision score.

Figure 5.5.2 Human-coded sentiment distribution by data source

Review-bias flag. Precision, recall, and F1 remained 1.00 against the three human-coded positive cases, all still in the original AB observations.

6. Discussion

The core pattern from the original study, communication support as the most reliable strength and price as the most reliable weakness, held up under two independent rounds of expansion across six different platforms, which is meaningful evidence that the pattern is not an artifact of the original nine clinics' selection. The Wonjin price/anesthesia complaint, the original Girin rushed-consultation-and-changing-price review, and Bbae/Teia/Dream/Lala's positive price narratives together describe the same underlying axis, price transparency and value, from both directions and across five different platforms.

The three-way batch comparison (Section 5.3) is this revision's central methodological contribution. It moves the paper from a general, literature-supported claim, that review platforms differ systematically[28], to a specific, dataset-internal demonstration: the same research team, asking the same five-dimension questions of a similar number of similar clinics, recovered a materially different sentiment mix depending only on which platform they looked at. A researcher relying solely on Google Maps for this kind of study would conclude that international patient experience contains essentially no mixed-sentiment accounts; a researcher who deliberately diversified sourcing would conclude the opposite. Neither conclusion is "wrong" given its own data; the point is that the platform choice itself is doing analytical work that a single-platform methodology section would not surface.

The Kai/Woo-a "no unnecessary upselling" pattern found in the non-Google batch is a useful complement to the price-complaint pattern found elsewhere. Read together, they suggest the axis patients care about is not price level in isolation but whether the recommendation they received was calibrated to their actual need rather than to a maximized sale, consistent with the SERVQUAL "reliability" framing already applied to this dataset in the prior revision[11]: what erodes trust is not a high price but a mismatch between what was promised or implied and what was delivered.

On the methodological question at the center of RQ4, the review_credibility keyword-collision finding (Section 5.5) is a small but genuine contribution: it is a directly observed instance, not just a cited literature claim, of an AI classifier's apparent competence depending on incidental surface features of the specific corpus it is run against[30]. Because the classifier was deliberately left untuned after this was discovered, the metrics in this paper are a more honest, if slightly less flattering, picture of what a simple, transparent lexicon classifier actually does than a version massaged after seeing the evaluation set would have produced.

7. Limitations

Several limitations bear on how the findings above should be read, extending those already documented for the original and round-one datasets.

The minimum-coverage rule (Section 4.2) improves comparability but changes what a clinic's profile means: a clinic with exactly three populated dimensions (for example Grand, Woo-a, Kai, Teia, or JW) is not directly comparable to one with all five populated (AB, Pitangui, VG), and the per-clinic "lowest dimension" figures in Section 5.1 should be read with this asymmetry in mind.

Translation and summarization of the Korean-language sources (Gangnam Unni, Babitalk, Sungyesa) were performed by the authors alone, without independent back-translation or a second bilingual reviewer checking the paraphrase against the original text. This is a genuine methodological gap relative to the rigor a fully resourced study would apply, and it is disclosed rather than hidden: a systematic misreading of tone or emphasis in the Korean-language observations is possible and has not been ruled out.

The non-Google batch is also still small, 8 new clinics and 18 observations, collected in a single research session on 2026-08-15. It is not a random or complete sample of any of the six platforms it draws from, any more than the Google-sourced batch was a random sample of Google. The three-way comparison in Section 5.3 should be read as evidence that platform does matter, not as a precise estimate of how much it matters or a claim that non-Google platforms are more representative in some absolute sense. They simply differ from Google, and from each other, in ways this dataset is now positioned to describe rather than assume.

The Grand Plastic Surgery observation (Section 5.4) also required judgment calls, about reporting a serious, corroborated but historical allegation; a different research team might have handled it differently. The two WhatClinic reviews added to backfill Grand's dimension coverage are unrelated in content to the historical controversy and should not be read as addressing or resolving it either way.

All limitations documented for the original 9-clinic sample and the round-one expansion continue to apply where relevant: the sample remains small and non-random; AI classifications remain a research aid subject to error and are reported here as a comparison against human coding, not as an independent or authoritative source of findings.

8. Ethical Considerations

This study concerns patients' reported medical experiences and treats their privacy accordingly across all three data-collection passes. The entire dataset was compiled through desk research on material patients and clinics had already made publicly available on the internet; no patient or clinic was contacted, interviewed, or surveyed at any point, and no consent process was required or applicable for that reason. No personally identifiable patient information is included in this paper; every narrative was paraphrased and anonymized (no reviewer names, photos, or profile detail), consistent across all three passes. All internet-based data collection (Google Maps, WhatClinic, Gangnam Unni, Babitalk, Sungyesa, Bookimed, Threads) was performed by manually reading public listings, not by automated bulk scraping, specifically to avoid the terms-of-service and data-protection concerns bulk collection would raise; the number of clinics and reviews involved was kept small for the same reason, and every observation records its exact platform and collection date for traceability. The study does not evaluate or claim to evaluate surgical safety, medical competence, or clinical outcomes, and it does not rank or recommend individual clinics; ratings and narratives are treated throughout as evidence of reported experience, not as objective measures of medical quality. Where this paper discusses a serious historical allegation against a real, named clinic (Section 5.4), it does so only after independent corroboration, with that corroboration and its historical scope explicitly stated.

One illustrative figure in this paper (a labeled before/after facial-contouring graphic used to orient a reader unfamiliar with the procedures discussed) is a generic informational/marketing-style graphic, not a photograph of any patient in this study's dataset or associated with any specific clinic in it; it is captioned accordingly and is not presented as evidence about any dataset clinic's outcomes.

9. Conclusion

This revised study extended a preliminary, mixed-methods analysis of international patient experience in South Korean plastic surgery from 9 clinics and 36 hand-compiled observations to 23 clinics and 74 observations, collected across three distinct rounds and seven different sources (the original document, Google Maps, WhatClinic, Gangnam Unni, Babitalk, Sungyesa, and Bookimed, plus one Threads post). The core substantive finding held up across all of them: communication support, particularly translation and interpreter services, remained the most reliably praised part of the patient experience, and price and value transparency remained the most reliably weak dimension, though the expansion also surfaced the dataset's clearest positive price narratives, showing that dimension is contested rather than uniformly negative. The expansion's more important contribution is methodological: a deliberate, two-round, six-platform data-collection design produced direct, measured evidence that the platform a review is sourced from changes what a review dataset can say, independent of the underlying clinics' actual quality. The Google-sourced batch contained no mixed-sentiment accounts at all, while diversifying sourcing recovered them. On the AI-assisted classification question central to RQ4, accuracy held roughly steady (82% sentiment accuracy, macro theme F1 of 0.74) while a specific, undisguised classifier failure mode, a keyword collision between "review" and "reviewer," was identified and reported rather than quietly fixed, in the interest of giving an honest account of what a simple, transparent classifier actually does. The next stage of this research is to expand data collection under a standardized, pre-registered protocol across all six platforms identified here, add independent verification for the Korean-language translations, and report a larger-sample precision/recall/F1 evaluation than this preliminary, three-batch sample can fully support.

References

  1. Rastogi, I., & Puri, P. (2026a). Understanding international patient experiences in South Korean plastic surgery: An AI-assisted analysis of patient reviews, service quality, communication, and trust [Unpublished manuscript]. Department of Computer Science Engineering, Amity School of Engineering & Technology, Amity University Uttar Pradesh.
  2. Korea Times, The. (2016, July 24). Samsung Medical Center haunted by 'ghost surgery'. The Korea Times.
  3. Bookimed. (2026). Dr. Im Young Min: prices, reviews (VG Plastic Surgery, Seoul). Retrieved August 15, 2026.
  4. Holliday, R., Bell, D., Cheung, O., Jones, M., & Probyn, E. (2017). Trading Faces: the 'Korean Look' and Medical Nationalism in South Korean Cosmetic Surgery Tourism. Asia Pacific Viewpoint. https://doi.org/10.1111/apv.12154
  5. Koh, C. (2017). Characteristics of cosmetic medical tourism in Korea.
  6. Park, J., et al. (2021). A Look at Collaborative Service Provision: Case for Cosmetic Surgery Medical Tourism at Korea for Chinese Patients. International Journal of Environmental Research and Public Health, 18(24), 13329. https://doi.org/10.3390/ijerph182413329
  7. Yom, Y., et al. (2015). Factors Influencing Chinese Customers' Loyalty to Korean Medical and Tourism Services. Journal of Korean Academy of Nursing Administration.
  8. Anvarjonov, U. N. B., & Um, K.-H. (2023). The Effect of Social Media Marketing Capability on International Patient Satisfaction through Perceived Risk in the Medical Tourism Context. Journal of Korean Society for Quality Management, 51(2), 203-221. https://doi.org/10.7469/JKSQM.2023.51.2.203
  9. Rastogi, I., & Puri, P. (2026b). K-PSI research proposal: Understanding international patient experiences in South Korean plastic surgery [Unpublished research proposal]. Department of Computer Science Engineering, Amity School of Engineering & Technology, Amity University Uttar Pradesh.
  10. Liu Rui, et al. (2016). A Study on the Korean Medical Service of the Foreign Patients Patronage: Based on the Case of Chinese Cosmetic Surgery Patients in Korea. Korea International Trade Research Institute.
  11. Jonkisz, A., Karniej, P., & Krasowska, D. (2022). The Servqual Method as an Assessment Tool of the Quality of Medical Services in Selected Asian Countries. International Journal of Environmental Research and Public Health.
  12. Guiry, M., & Vequist, D. G. (2011). Traveling Abroad for Medical Care: U.S. Medical Tourists' Expectations and Perceptions of Service Quality. Health Marketing Quarterly, 28(3), 253-269.
  13. Qolipour, M., et al. (2018). Assessing Medical Tourism Services Quality Using SERVQUAL Model: A Patient's Perspective. Iranian Journal of Public Health.
  14. Alruwaili, K., et al. (2026). Why Do International Patients Come Back? An S-O-R Explanation of Service Quality, Trust, and Satisfaction in Saudi Hospitals. Veredas do Direito.
  15. Tang, J. E., et al. (2022a). How Are Patients Reviewing Spine Surgeons Online? A Sentiment Analysis of Physician Review Website Written Comments. Global Spine Journal.
  16. Langerhuizen, D., et al. (2020). Analysis of Online Reviews of Orthopaedic Surgeons and Orthopaedic Practices Using Natural Language Processing. The Journal of the American Academy of Orthopaedic Surgeons.
  17. Cho, L. D., et al. (2022). Sentiment Analysis of Online Patient-Written Reviews of Vascular Surgeons. Annals of Vascular Surgery.
  18. Butler, L. R., et al. (2022). Building better pediatric surgeons: A sentiment analysis of online physician review websites. Journal of Children's Orthopaedics.
  19. Tang, J. E., et al. (2021). Using Sentiment Analysis to Understand What Patients Are Saying About Hand Surgeons Online. HAND.
  20. Vought, V., et al. (2024). Application of sentiment and word frequency analysis of physician review sites to evaluate refractive surgery care. Advances in Ophthalmology Practice and Research.
  21. Assessing Patient Satisfaction Following Blepharoplasty Using Social Media Reviews. (2022). Aesthetic Surgery Journal, 42(3), NP179-NP185.
  22. Ramnarine, A. K. (2023). Unsupervised Sentiment Analysis of Plastic Surgery Social Media Posts. arXiv:2307.02640.
  23. Alaniz, L., et al. (2025). Patient Satisfaction Following Top Surgery: A RealSelf Analysis Using Advanced Natural Language Processing. Plastic and Reconstructive Surgery Global Open.
  24. Marinescu, I., Klein, N., Chamberlain, A., & Smart, M. (2021). Incentives can reduce bias in online employer reviews. Journal of Experimental Psychology: Applied.
  25. Karaman, H. (2020). Online Review Solicitations Reduce Extremity Bias in Online Review Distributions and Increase Their Representativeness. Management Science.
  26. Ai, J., Zhu, J., & Liu, Y. (2022). Effects of offering incentives for reviews on trust: Role of review quality and incentive source. International Journal of Hospitality Management.
  27. Liang, W.-Y., et al. (2025). The impact of mandatory disclosure on rewarding online reviews based on S-O-R theory. Asia Pacific Journal of Marketing and Logistics.
  28. Kordzadeh, N. (2019). Investigating bias in the online physician reviews published on healthcare organizations' websites. Decision Support Systems.
  29. Qahri-Saremi, H., Turel, O., & Ophir, Y. (2022). Negativity bias in the diagnosticity of online review content: the effects of consumers' prior experience and need for cognition. European Journal of Information Systems.
  30. Hill, C., et al. (2025). Large language models for thematic analysis in healthcare research: A blinded mixed-methods comparison with human analysts. PLOS Digital Health.
  31. Kornblith, A., et al. (2025). Analyzing patient perspectives with large language models: a cross-sectional study of sentiment and thematic classification on exception from informed consent. Scientific Reports.
  32. Bojić, L., et al. (2025). Comparing large language models and human annotators in latent content analysis of sentiment, political leaning, emotional intensity and sarcasm. Scientific Reports.
  33. plasticsurgery list.pdf. (n.d.). Preliminary hand-compiled research document supplied by the K-PSI research team, covering DA, Girin, View, Oneul, AB, Chiu, Braun, I-one, and HighEnd Plastic Surgery.
  34. Google Maps. (2026). Business listings and reviews for ID Hospital Korea, JW Plastic Surgery Clinic, Banobagi Plastic Surgery, Wonjin Plastic Surgery, Grand Plastic Surgery Korea, and Dream Plastic Surgery Clinic. Retrieved August 15, 2026.
  35. WhatClinic. (2026). Clinic profiles and verified patient reviews for Grand Plastic Surgery, Pitangui Medical & Beauty, Hyundai Mihak Plastic Surgery, and Bbae Clinic. Retrieved August 15, 2026.
  36. Gangnam Unni (강남언니). (2026). Community posts referencing Woo-a Plastic Surgery, Kai Plastic Surgery, and Teia Clinic. Retrieved August 15, 2026.
  37. Babitalk (바비톡). (2026). Community posts and hospital listing referencing Lala Plastic Surgery. Retrieved August 15, 2026.
  38. Sungyesa (성예사). (2026). Site-visit review referencing Lala Plastic Surgery. Retrieved August 15, 2026.
  39. Threads. (2026). Public post referencing Lala Plastic Surgery. Retrieved August 15, 2026.
  40. Duke University Press. (2022). Between Plastic Surgery and the Photographic Image. positions: asia critique.

0 comments

Sign in to join the discussion