Stigma, Concealment, and Marriage Matching: Biomarker Evidence from Korea
Ye Yuan
Abstract
Social norms and stigma are central to many economic outcomes but rarely observed directly. This paper provides a behavioural measure of stigma in a setting where the underlying trait can be biologically verified. Using twelve waves of the Korea National Health and Nutrition Examination Survey (2008-2021), I compare self-reported smoking against biomarkerverified smoking measured by urinary cotinine concentration for 8,768 married couples and 7,940 single adults. I document a 34 percentage point gendered concealment gap where biomarkerpositive women are far more likely than biomarker-positive men to self-report as nonsmokers, with married women showing an additional increment. Within the biomarker-positive sample of wives, concealment is systematically associated with husband education, surviving controls for the wife's own education and age. The choice of measurement affects empirical conclusions about marital sorting. The wife smoking-husband education relationship is 39 percent larger in magnitude under self-report than under biomarker measurement, and structural Choo-Siow surplus estimates differ systematically across measurement schemes. The gendered concealment gap has declined substantially over the study period.
Review reports
XSci (AI review)
Referee Report
Manuscript: "Stigma, Concealment, and Marriage Matching: Biomarker Evidence from Korea" Recommendation: Major revision
Summary and Recommendation
This paper asks whether social stigma can be measured through the concealment it induces, and whether stigma-driven misreporting distorts what applied researchers conclude about marriage matching. The setting is South Korean female smoking, a behaviour subject to strong gender-asymmetric social penalties and, unusually, one that can be biologically verified. Using twelve waves of the Korea National Health and Nutrition Examination Survey covering 2008–2011 and 2014–2021, the author compares self-reported current smoking against urinary cotinine, classifying respondents at or above 50 ng/mL as biological smokers and, among these, treating those who deny smoking as concealers. Four results are reported: a gender concealment gap of roughly 34 percentage points among biomarker-positive adults, with an additional 9.8 point increment for married women (Table 4); a positive association between husband education and wife concealment within the 835 biomarker-positive wives, at 1.6 percentage points per year of schooling and 1.3 points conditional on the wife's own education (Table 6); a divergence of approximately 39 percent between self-report and biomarker estimates of the wife-smoking education gradient, at −0.744 against −0.535 (Table 7); and a decline in female concealment from approximately 71 percent in 2008 to 46 percent in 2021.
The central idea is a good one, and the paper deserves to be developed rather than abandoned. Using a biomarker not merely to correct measurement error but to study the social forces that generate it is a genuinely novel move, the Korean setting is close to ideal for executing it, and linking the exercise to the bidimensional matching literature gives it a concrete payoff for applied readers. The robustness agenda in Section 4.6 anticipates most of the obvious objections, and the author is commendably explicit that no causal claims are being made. Major revision is nonetheless required. The manuscript does not yet establish that the concealment indicator measures what it is said to measure, since the biomarker admits false positives that are correlated with the paper's key regressors and the self-report definition imposes a screen that binds for women and not for men. The sample definitions underlying the headline statistics do not reconcile across tables, and the correlations that constitute the paper's novel contribution depend on specification choices whose consequences are not explored. And the matching application, which is where the paper's methodological claim lives, rests on a structural exercise that appears to be run on two different populations and on a reduced-form comparison that is never developed econometrically. One further matter of positioning deserves mention at the outset: the gendered concealment gap presented as the first of four findings is, on the manuscript's own account in Section 2.1, already established in the Korean literature, where Park et al. (2014) report under-reporting rates of roughly 60 percent for women against 13 percent for men on the same data source and at the same threshold. Acknowledging this, and leading instead with the marital dimension and the measurement lesson, would make the paper's contribution clearer rather than smaller.
Major Concern Category 1: What the Concealment Indicator Measures
1.1 Urinary cotinine as ground truth and the environmental exposure channel
The design rests on the premise, stated in Section 1, that the biomarker "provides objective ground truth that is unaffected by reporting incentives," and while the concluding section acknowledges that second-hand smoke may generate misclassification, the acknowledgement is not carried into the analysis. The difficulty is that exposure is not orthogonal to the outcomes of interest: Table A.2 reports that 74.4 percent of concealers' husbands are cotinine-positive smokers against 38.8 percent for true nonsmokers' husbands, so a nonsmoking wife of a heavy smoker is precisely the person most likely to be misclassified as a concealer. If even a modest share of the 473 women in that category are nonsmokers with high household exposure, the coefficients in Table 5 columns (5) and (6) partly reflect a mechanical exposure channel, and the contamination propagates into the husband-education results. The evidence offered against this reading is not yet sufficient, since the mean cotinine of 639.5 ng/mL cited in Section 5.6 is a weak statistic for a distribution the paper itself describes as sharply bimodal, and the restriction in Table A.5 column (3) to the 50–300 ng/mL window retains 251 of 835 biomarker-positive wives, suggesting that roughly three in ten sit in the lower intensity range. Reporting the full distribution of cotinine among concealers, re-estimating among wives of nonsmoking husbands, conditioning on the second-hand exposure items KNHANES collects, and showing sensitivity to higher thresholds would each address this directly.
1.2 The lifetime-cigarette screen binds asymmetrically by gender
Section 3.3 defines a self-reported current smoker as someone who has smoked at least 100 lifetime cigarettes and currently smokes daily or occasionally, following Chiappori et al. (2018). Because the condition is conjunctive, a woman below the lifetime threshold is classified as a nonsmoker however truthfully she answers, and the paper's own summary statistics show this screen binding for women alone. Table 1 reports that 8.3 percent of married women have smoked 100 lifetime cigarettes while 9.6 percent are cotinine-positive, so even if every woman passing the threshold were also biomarker-positive, at least 115 of the roughly 847 biomarker-positive married women could not be classified as current smokers under any answer they might give; set against 473 concealers, this is a lower bound of perhaps a fifth of the category. The same inversion appears among single women (19.1 against 19.5 percent), while for married and single men the screen is far from binding (80.5 against 42.3, and 65.0 against 54.8). The asymmetry runs in exactly the direction that produces the headline result, and is reinforced by the intensity margin, since married women report 8.628 cigarettes per day against 15.639 for married men and a mean initiation age of 23.995 with a standard deviation of 9.280. The gap between 45.76 and 10.70 percent in Table 3 is far too large to be wholly definitional, but the surviving magnitude is the quantity of interest, and reporting the gap under a definition that drops the lifetime screen would isolate it.
1.3 The setting in which disclosure is observed
What the data record is an answer given to a KNHANES enumerator, and the manuscript does not establish that this reflects an internalised social cost rather than the circumstances of the interview. Section 3.1 notes that the survey combines a face-to-face interview, a self-administered questionnaire, and an examination, but never states which instrument carries the smoking items, whether answers are spoken or written, or whether other household members may be present. This matters because the situational reading is a direct competitor that would generate the paper's signature correlation through a channel unrelated to norms: if the items are administered face to face at home, whether the husband is present becomes a determinant of the wife's answer, and presence during a daytime visit is strongly related to employment, occupation, and hours, which are in turn related to education. Documenting the instrument and mode would settle the first-order question, and if the items are face to face, conditioning on husband employment status and hours would provide a test. A related pattern in the paper's own results bears on the mechanism and is currently passed over: if concealment served to protect a marriage, relative spousal position should matter, yet Table 6 shows neither the education gap (0.004, standard error 0.005) nor the marry-up indicator (0.018, standard error 0.036) associated with concealment while the husband's absolute education level is.
1.4 The absence of a framework linking concealment to a stigma parameter
The abstract describes the paper as providing "a behavioural measure of stigma," yet no model connects the measured quantity to a stigma cost, and the concluding section itself concedes that "the object measured here is not stigma directly, but stigma-induced concealment." The gap is consequential because the observed concealment rate is a composite of at least four things: the social cost attached to the trait, the respondent's subjective probability that a survey answer will reach anyone who might impose that cost, the salience and self-perception of the behaviour, and, through the threshold, smoking intensity. Without separating these, the cross-gender comparison cannot distinguish a larger penalty from a different perceived disclosure risk or a different self-concept about light, intermittent smoking, and the comparative statics across marital status and time cannot be signed. The consequences are concrete: the prediction in Section 4.2 that the stigma hypothesis implies a smaller concealer than discloser coefficient is asserted rather than derived and is close to definitional; the marital increment is read as intensified stigma when marriage plausibly changes the audience, the cost, and the interview setting at once; and the declining trend is read as falling stigma when rising confidence in confidentiality would produce the same pattern. A compact two-page framework in which an agent trades a disclosure cost against a reporting cost would generate testable comparative statics and would also clarify which estimand the measurement comparison in Section 5.4 should target.
Major Concern Category 2: Sample Construction and the Identification of the Concealment Correlations
2.1 Inconsistent sample definitions behind the headline statistics
Several closely related quantities do not reconcile, and the discrepancies reach the abstract. Table 3's counts of 8,050 biomarker-positive married men and 2,275 married women cannot come from the linked-couples sample, where 8,822 married men are 42.3 percent cotinine-positive; they correspond instead to Appendix Table A.1, where 16,989 married men at a positivity rate of 0.474 yields approximately 8,053. But A.1's 27,550 married men and 37,969 married women sum with the never-married counts to exactly the 79,208-person base sample, implying that "married" there denotes all ever-married respondents, which would also explain the implausible excess of women. If so, the abstract's statement that roughly 46 percent of biomarker-positive wives report as nonsmokers describes ever-married women, whereas the couples sample gives 473 of 835, or 56.6 percent — the figure that appears in Table 1 under a note stating, incorrectly, that the denominator includes nonsmokers. Compounding this, Section 3.2 reports an analytic sample of 8,768 couples while Sections 3.6 and 3.7 and Tables 1, 2 and 7 use 8,822; the biomarker-positive wife subsample is 835 in Tables 6 and A.5 but 848 in Tables A.4 and A.7; and the three groups sum to exactly 8,768 although Section 3.4 states that a fourth category is excluded. Since the contribution is a measurement contribution, a reader must be able to trace each figure to one consistently applied population.
2.2 Selection into the linked-couples sample
Two filters shape the couples sample and neither is characterised. The linkage rule retains only households with exactly one respondent–spouse pair, excluding 4,931 married adults whom Section 3.2 identifies as older coresidents of multi-generational households. This is an awkward exclusion for this paper in particular, since Section 2.2 grounds the entire interpretation in Confucian family organisation and intra-household hierarchy, and a wife living with her husband's parents plausibly faces denser social monitoring than one who does not. The second filter is larger in volume: only 9,509 of 22,778 linked couples have cotinine on both spouses, a retention rate of 41.7 percent, and Section 3.1 describes cotinine as measured for "a rotating subsample" without stating whether the rotation operates at the individual or household level — a distinction that determines whether requiring both spouses is innocuous or severely selective. Participation in the examination conditional on selection is a further margin, plausibly correlated with the health behaviour under study. The validation offered in Section 3.6 cannot do this work, since it compares the couples sample against an ever-married benchmark, and the differences it does show are not small: married women average 7.825 years of schooling in Table A.1 against 8.945 in Table 1, with cotinine positivity of 0.116 against 0.096. The informative comparison, between linked couples with and without both-spouse cotinine, is available and should be reported.
2.3 Age and cohort confounding in the three-group comparison
The central descriptive result of Section 5.2 is that concealers occupy an intermediate position between true nonsmokers and disclosers on husband characteristics, but this ordering exists only after conditioning and reverses the raw pattern. Table A.2 shows concealers' husbands with 10.579 years of schooling against 10.192 for true nonsmokers' husbands and 9.710 for disclosers', making concealers' husbands the most educated group unconditionally; the conditional estimates of −0.367 and −0.753 in Table 5 are produced by the controls, as Section 5.2 concedes when it attributes the raw ordering to age differences. The concern is that the age differences are large and the adjustment restrictive: concealers average 45.8 years against 51.7 for true nonsmokers, with husbands at 48.7 and 54.9, and over the cohorts spanned here Korean educational attainment expanded rapidly and nonlinearly, a gradient Table 1 illustrates with married women at 8.945 years against 12.970 for single women averaging 27.0 years of age. When the grouping variable is strongly associated with age and the outcome is strongly and nonlinearly associated with age, a linear control is unlikely to absorb the confound, and a sign reversal between raw and adjusted comparisons is the signature of residual confounding. Birth cohort is already constructed and used elsewhere; re-estimating Tables 5 and 6 with cohort fixed effects or single-year-of-age indicators, or comparing within narrow age bands, would settle whether the ordering survives.
2.4 The functional form of the wife-education control
The abstract's claim that concealment is associated with husband education "surviving controls for the wife's own education" rests on Table 6 column (5), where the husband coefficient falls from 0.016 to 0.013 and remains significant while wife education in years (0.008, standard error 0.006) does not. The appendix tells a different story when the wife's education enters as a binary indicator: Table A.4 reports wife high-education coefficients of 0.185 and 0.169, and Table A.7 reports 0.183 and 0.164, all highly significant on the same biomarker-positive wife sample and comparable in magnitude to the husband high-education coefficient of 0.145 in Table 6 column (2). Meanwhile wife education in years is small and insignificant in Table A.7 column (3) as well (0.005, standard error 0.007). Because husband and wife education are strongly associated — Table 7 reports a spousal homogamy coefficient of 0.624 — the two are close substitutes in a regression, and how much of the common variation is attributed to the husband depends on how the wife's education is parameterised. If the relationship is concentrated at the high-school margin, as the binary results suggest, a linear control will not absorb it. The decisive specification, with the wife's high-education indicator or a full set of her education categories as the control, is never reported and should be.
2.5 The asserted equivalence of the two framings
Section 5.3 states that the reverse-framing regression "is descriptively equivalent to the three-group comparison in Section 5.2" and offers a specific reconciliation, claiming that the implied concealer-versus-discloser contrast of +0.386 years of husband education "matches" the coefficient of 0.016 in Table 6 column (1). No derivation is supplied, and the two quantities are not obviously commensurable, since the algebra connecting a reverse regression to a mean difference runs through the residual variance of husband education, which is never reported; using the unconditional variance implied by Table 1 gives an implied slope around 0.004, and reconciliation would require the residual variance to fall to roughly a third of its unconditional value. That is possible but is an empirical claim the paper neither makes nor supports. There is also a substantive difficulty: the two regressions share neither sample (8,768 versus 835), baseline (true nonsmokers versus disclosers), nor control set, so describing them as "the same fact from different perspectives" overstates the case, and Table A.5 column (2), which estimates the concealer coefficient on husband education within the 835, is the genuinely comparable specification. Relatedly, the narrative slides between two findings of opposite sign — concealers' husbands less educated than true nonsmokers' in Section 5.2, higher husband education predicting concealment in Section 5.3 — without flagging the change in comparison group, and the abstract reports only the second.
Major Concern Category 3: The Matching Application and the Discipline of Inference
3.1 The econometrics of the measurement comparison
The paper's methodological claim rests on Table 7, where the wife-smoking gradient in husband education is −0.744 under self-report and −0.535 under biomarker measurement, a divergence attributed to self-reported female smokers being "a small, negatively selected subset," while the near-identity of the homogamy coefficients (0.240 and 0.235) is attributed to both spouses' smoking being "mismeasured in parallel." Neither explanation is developed, and both appear to sit awkwardly with what a characterisation of the misclassification process would predict. The error here is one-sided by construction, since a woman who self-reports as a smoker is biomarker-positive, so there are essentially only false negatives, at a rate of 0.458 among biomarker-positive married women. Under one-sided misclassification independent of the outcome, a rare binary regressor attenuates only slightly, because the misclassified positives are diluted in a large negative group: with a true positive rate of 0.096 and a false-negative rate of 0.458, the implied attenuation factor is roughly 0.95, which would predict a self-report gradient near −0.51 rather than one 39 percent larger. This is a stronger conclusion than the paper draws, since it implies the entire divergence is differential selection rather than attenuation, and it also sharpens the external lesson: where concealment is unselective, self-report would give nearly the right answer. The homogamy result is correspondingly harder to explain, since the two false-negative rates of 0.458 and 0.107 are not parallel and the same logic predicts a self-report coefficient near 0.20 rather than 0.240. A short subsection deriving both probability limits under a stated misclassification process, supported by calibration, would convert an intuition into a result.
3.2 The Choo-Siow surplus exercise
Three features of the structural complement require attention. Most seriously, the two panels of Table A.8 do not appear to describe the same couples: the bracketed cell counts in Panel A sum to 19,013 while those in Panel B sum to 8,822, the latter matching the analytic sample in Table 2. If self-report is being computed on all linked couples while cotinine is restricted to those with both-spouse measurements, the comparison confounds a change in measurement with a change in who is included, and the conclusion that self-report "systematically distorts the inferred matching surplus pattern" cannot be sustained; restricting both panels to a common set of couples would remedy this. Second, since the surplus estimator is a deterministic function of cell counts — twice the log of the couple count less the logs of the two singles counts — reclassifying concealers as smokers necessarily raises smoker diagonals and lowers nonsmoker diagonals, and the observed movement on the high-education smoker diagonal from −2.79 to −1.89 is close to the 1.03 that the count change alone implies. Presenting this as an economic finding risks describing the arithmetic of the estimator, and a benchmark — the surplus matrix under random reclassification — would show whether anything beyond arithmetic is happening. Third, the singles pool supplying the unmatched counts averages 29.4 and 27.0 years of age against 54.3 and 51.1 for the matched couples, and twelve waves and seventeen regions are pooled into a single undifferentiated market, so the level of the estimated surplus carries no structural interpretation as currently constructed.
3.3 Smoking as a matching trait
Sections 4.2 and 6 acknowledge that smoking is observed at survey date rather than at marriage formation, but treat this as a caveat on causal language when it is in fact a constraint on what the matching estimates can mean. Table 1 reports that married women's mean age at first smoking is 23.995 with a standard deviation of 9.280, an unusually wide distribution implying that a considerable minority report initiation at ages above the typical age at first marriage for Korean women over this period; for those women smoking was not a trait available to sort on, and a regression of husband education on current wife smoking measures the correlation between post-marital behaviour and spousal characteristics. The contrast with married men, at 19.757 with a standard deviation of 4.240, makes the asymmetry concrete: male smoking is largely pre-marital in this sample while female smoking substantially is not, and because the paper's central result is an asymmetry between the wife and husband directions of the matching regressions, this initiation asymmetry is a competing explanation running the same way. Conditioning on reported initiation age, or restricting to wives whose initiation precedes a plausible marriage age, would test it. A second issue is that the sample is a stock of surviving marriages pooled across twelve waves, so the estimated patterns reflect both sorting at formation and differential dissolution afterward — a point with particular force for the Choo-Siow exercise, which treats the observed distribution as an equilibrium matching.
3.4 The time trend and the cohort-period decomposition
The fourth headline finding is read as evidence of changing norms, but four coincident changes are not addressed. The laboratory method changed from gas to liquid chromatography in 2016, and validation for accuracy does not establish comparability of a binary classification at a fixed threshold; since a substantial share of concealers sit in the lower intensity range, even modest differences in calibration or recovery would shift the measured rate for technical reasons. The questionnaire coding changed in 2010, when the current-smoking item began distinguishing daily from occasional smoking, and offering an explicit intermediate category would plausibly raise measured disclosure thereafter — which matters because the two earliest waves anchor the high end of the reported decline. The denominator's composition is itself shifting as female smoking prevalence and the age structure of the married population evolve, with roughly seventy biomarker-positive wives per wave supporting the 2008 endpoint of 71 percent. Finally, the cohort-and-period specification in equation (7) is not identified, since Section 3.5 defines cohort as wave year minus age, so including both sets of fixed effects absorbs rather than estimates the age effect — as Table A.4 column (3) confirms by reporting no age coefficient. The conclusion that "both period effects and cohort replacement contribute" therefore requires an identifying restriction that is neither imposed nor discussed. Plotting the cotinine distribution by wave, estimating separately before and after 2016, and either defending a normalisation or presenting cohort profiles descriptively would address these.
3.5 The reading of null and adverse results
A paper whose contribution is empirical is judged on the discipline of its inference, and several passages fall short of the standard the rest of the manuscript sets. Table A.5 column (3) restricts to cotinine between 50 and 300 ng/mL, leaving 251 observations and yielding −0.189 with a standard error of 0.587 — opposite in sign to the main estimate of +0.456 and with an interval comfortably containing it. Section 5.6 reads this as both a power limitation and as evidence that "enhances our main specification results," which are mutually inconsistent: a test that cannot reject anything cannot corroborate anything, and the honest statement is that the intensity concern remains open. Similarly, the male placebo's wife-education coefficient of +0.179 with a standard error of 0.150 spans an interval that does not exclude effects of the magnitude found on the female side, yet four imprecise nulls are treated as a demonstrated absence, while the one significant result (−0.061) is explained away without a test. The non-monotone cohort interaction pattern in Table A.7 column (3), estimated from 835 observations across five cohorts, carries a post-hoc narrative about "a unique combination of strong residual stigma and active marriage-market opportunities" for which no independent evidence is offered; the same column's cohort main effects are evaluated at husband education of zero, far outside a support centred near 10.177 years, and differ in sign from column (1) without the text alerting the reader. Across Tables 5, 6, A.6 and A.7 a large number of hypotheses is examined without any adjustment for multiplicity.
Minor Issues
-
Attribution of the inverse hyperbolic sine transform. Section 4 states that the author follows Chen and Roth (2024) in using the asinh transform, which inverts their argument: they show that estimands from log-like transformations, explicitly including asinh, depend on the units of the outcome, and recommend alternatives. Relatedly, the asinh cotinine coefficient of −0.008 with a standard error of 0.101 in Table A.5 is described as small and insignificant on a scale whose units are not interpretable; reporting how the concealer coefficient moves with and without the control would be more informative.
-
Mechanically redundant columns. Because the education gap is husband's schooling minus wife's and wife education is a control, Table 5 column (3) reproduces column (1) exactly (−0.367 and −0.753, with standard errors 0.138 and 0.175); the same holds between columns (1) and (3) of Table A.6. Presenting these as independent corroboration overstates the evidence, and the marry-up result in column (4) is described as showing "the same patterns" although the concealer coefficient (−0.037, standard error 0.023) is not significant.
-
No test of the measurement difference, and non-common samples. The comparisons in Table 7 are made by inspecting point estimates on overlapping samples whose estimates are correlated, so the headline "39 percent" is a ratio reported without any measure of precision. Stacking the specifications with couple-clustered standard errors would deliver both a test and an interval. Observation counts also vary across columns (8,743; 8,822; 8,785; 8,822; 8,777; 8,822), and across Table 5 as well, so sample composition is not held fixed in a comparison specifically about measurement.
-
Survey design and estimator choice. KNHANES is a stratified multi-stage probability sample of households, yet no analysis uses sampling weights and standard errors are heteroskedasticity-robust rather than clustered, although spouses are observed jointly within households. Separately, all concealment regressions are linear probability models, with the husband specifications sitting near an outcome mean of 0.11; a logit or probit check would be reassuring, and no standard errors are reported for the Choo-Siow estimates despite cells as small as 17 couples.
-
Overinterpreted nulls. Section 5.6 concludes from a metropolitan coefficient of −0.021 with a standard error of 0.035 that "female smoking stigma is uniform across Korean geography," although the interval spans roughly −9 to +5 percentage points on a base near 46 percent. The male placebo likewise contrasts significance across separately estimated models with 3,707 and 835 observations, which is not a test of coefficient equality; a pooled specification interacting concealer status with sex would provide one.
-
Numerical reconciliations. Table 4 column (4), a robustness check using a stricter definition, has more observations (13,563) than the baseline columns (13,232). Section 5.5 describes the decline as "roughly 25 percentage points" from the figure while the trend of −0.022 per year implies 28.6 over the period. The same section describes husband under-reporting as showing "no clear trend" and then as significantly declining at −0.003 per year. Each should be reconciled.
-
Definitions and terminology. Section 3.3 describes the daily-only indicator as "equivalent to the 2008–2009 definition," but footnote 1 shows that the earlier coding did not separate daily from occasional smoking, so the later restriction is stricter rather than equivalent. The text also moves between "concealment," "under-reporting," and "misreporting" across sections and tables, and uses "stigma" for what the survey literature would call socially desirable responding; since intent is the object of interest, consistent terminology with a stated definition would help.
-
Comparisons lacking common support, and framing. The female-by-married interaction in Table 4 column (3) compares married women averaging 51.1 years with single women averaging 27.0 under a linear age control, and should be shown to survive restriction to overlapping ages. The marital increment itself (45.76 against 39.51 percent raw) is modest relative to a gender gap roughly five times larger. The assertion in Section 5.1 that the gap exceeds "any previously documented gender concealment gap" is an unverifiable superlative that the careful Exley et al. comparison in the same paragraph renders unnecessary.
-
Specification and reporting gaps. The metropolitan indicator derives from the same seventeen-category variable as the region fixed effects specified in Section 4.6 and may be collinear with them; Table A.7 does not report which fixed effects are included. No table reports the wave composition of the couples sample despite breaks in the linkage variable in 2010 and the assay in 2016, and Figure 2 lacks confidence bands and per-wave cell sizes. No table reports dependent-variable means or fit statistics, which makes magnitudes hard to assess. The education mapping in Section 3.5 assumes completion and assigns a single value to graduate attainment, and sensitivity to an alternative mapping would be useful given that husband education is the key parameter.
-
Presentation. Several cross-references read "Table Table 1" and "Figure Figure 1b"; Table 6 lists "Husband Hige-edu"; the Ho et al. (2004) entry lacks a title and source; the Section 5 roadmap skips Section 5.3, which contains the regression underlying an abstract claim; and the single-authored manuscript alternates between "I" and "our." The introduction also contains constructions that impede comprehension ("this paper contribute," "The another strand," "the time trends of its revolution"), and Camerer et al. (2004) is characterised as a coordination-game elicitation method where Krupka and Weber (2013), already cited, is the apt reference. Finally, the JEL codes and keywords omit the measurement contribution and the assortative-mating literature, the two audiences the paper most wants to reach.
Missing and Inadequately Integrated Literature
Cited but Inadequately Integrated
Several works the manuscript cites contain precisely the analytical content its arguments require. Benowitz et al. (2019) is invoked for the general validity of cotinine, but the 2019 update discusses cutoff selection explicitly and emphasises that appropriate thresholds depend on the prevalence and intensity of second-hand exposure in the population studied — exactly the issue raised in Section 1.1 above, and the natural basis for justifying the 50 ng/mL threshold in a sample where most concealers' husbands are themselves biomarker-positive. Benowitz, Hukkanen and Jacob (2009), cited for pharmacology, documents inter-individual variation in nicotine metabolism including differences by sex, which would affect the probability of crossing a fixed threshold at given consumption and thus contribute to a measured gender gap independent of disclosure. Kweon et al. (2014) is cited once for the survey design but not used to establish whether the cotinine rotation operates at the individual or household level, a question that determines how selective the both-spouse requirement is.
On the econometric side, Celhay, Meyer and Mittag (2024, 2025) are the closest antecedents and are used only for framing, although they develop machinery for characterising non-classical measurement error and tracing it into regression estimates that would directly support the derivation recommended in Section 3.1. Chen and Roth (2024) is cited in a manner that reverses its message. Choo and Siow (2006) supplies the surplus formula but not its market structure, and the manuscript should engage with the framework's requirement that unmatched agents represent the outside option for the same market as the matched ones. Chiappori, Oreffice and Quintana-Domeque (2018) supplies the smoker definition including the lifetime screen, without noting that a definition calibrated to a U.S. population where it does not bind behaves differently where it does.
Substantively, Park et al. (2014) and Jung-Choi et al. (2011) establish the gendered discrepancy on the same data and should be treated as the benchmark the paper confirms and extends rather than as background to a new discovery; reconciling their prevalence figures with the manuscript's own would also strengthen confidence in the sample construction. Exley et al. (2024) is the closest comparison, but its proposed mechanism is strategic inference rather than norms, and spelling out what that mechanism would predict in the Korean setting — and whether the data can distinguish the two — would sharpen the contribution. Woo (2018) documents the strategies Korean women use to manage disclosure across settings, which speaks directly to whether concealment is audience-specific. Moffitt (1983) and Besley and Coate (1992), cited only as list entries, contain the modelling apparatus that Section 1.4 finds missing: both formalise stigma as a cost entering a participation decision and show how it can be identified from observed behaviour.
Truly Missing Papers
Four strands of work bear directly on the manuscript and are absent. On misclassification, Aigner (1973) provides the original treatment of a binary regressor measured with error and the attenuation result the paper's comparison implicitly invokes; Bollinger (1996) derives sharp bounds for regression parameters under binary misclassification and provides estimators for them, which is an unusually good fit here because the biomarker gives direct purchase on the misclassification probabilities; Hausman, Abrevaya and Scott-Morton (1998) addresses misclassification in a discrete dependent variable, which is the situation in the homogamy specification; Mahajan (2006) establishes identification when misclassification probabilities depend on covariates, exactly the manuscript's case since concealment varies with husband education; and Bound, Brown and Mathiowetz (2001) is the standard survey of validation studies against external benchmarks and supplies comparison magnitudes.
On matching, Galichon and Salanié (2022) generalises Choo and Siow under separability and, critically here, supplies estimation and inference procedures that would allow standard errors for the surplus estimates rather than leaving the structural results exploratory on that ground; Chiappori and Salanié (2016) sets out the identification requirements and market-definition conventions the exercise needs to satisfy; and Dupuy and Galichon (2014) offers tools for judging how many matching dimensions data can support, bearing on the four-type discretisation.
On stigma and disclosure, Goffman (1963) supplies the concept and specifically the distinction between discredited and discreditable traits, the latter being those that can be concealed — a distinction on which the entire design turns; Akerlof and Kranton (2000) provides the identity framework that Section 2.2 describes verbally; Bénabou and Tirole (2006) is the standard model of behaviour under social image concerns and the most direct template for the disclosure decision recommended above; Andreoni and Bernheim (2009) isolates audience effects experimentally, which bears on whether the measure reflects an internalised cost or a situational presentation; Bertrand, Kamenica and Pan (2015) documents how within-household gender norms shape both behaviour and reporting; Bharadwaj, Pai and Suziedelyte (2017) is the closest published analogue to this design, finding under-reporting of mental health conditions roughly 36 percent of the time against much lower rates for non-stigmatized conditions; and Tourangeau and Yan (2007) reviews how reporting of sensitive behaviours varies with interview mode, privacy and third-party presence.
On the data and temporal analysis, Heckman and Robb (1985) is the canonical statement of the age-period-cohort identification problem and is load-bearing for equation (7); Deaton (1997) provides the standard treatment of cohort profiles from repeated cross-sections; Solon, Haider and Wooldridge (2015) sets out when weighting complex survey data is appropriate; and Romano and Wolf (2005) provides procedures for controlling error rates across the many hypotheses examined here.
References
Aigner, D. J. 1973. Regression with a binary independent variable subject to errors of observation. Journal of Econometrics 1(1):49–59.
Akerlof, G. A., and R. E. Kranton. 2000. Economics and identity. Quarterly Journal of Economics 115(3):715–753.
Andreoni, J., and B. D. Bernheim. 2009. Social image and the 50-50 norm: A theoretical and experimental analysis of audience effects. Econometrica 77(5):1607–1636.
Bénabou, R., and J. Tirole. 2006. Incentives and prosocial behavior. American Economic Review 96(5):1652–1678.
Bertrand, M., E. Kamenica, and J. Pan. 2015. Gender identity and relative income within households. Quarterly Journal of Economics 130(2):571–614.
Bharadwaj, P., M. M. Pai, and A. Suziedelyte. 2017. Mental health stigma. Economics Letters 159:57–60.
Bollinger, C. R. 1996. Bounding mean regressions when a binary regressor is mismeasured. Journal of Econometrics 73(2):387–399.
Bound, J., C. Brown, and N. Mathiowetz. 2001. Measurement error in survey data. In J. J. Heckman and E. Leamer, eds., Handbook of Econometrics, Volume 5, Chapter 59. Amsterdam: Elsevier.
Chiappori, P.-A., and B. Salanié. 2016. The econometrics of matching models. Journal of Economic Literature 54(3):832–861.
Deaton, A. 1997. The Analysis of Household Surveys: A Microeconometric Approach to Development Policy. Baltimore: Johns Hopkins University Press.
Dupuy, A., and A. Galichon. 2014. Personality traits and the marriage market. Journal of Political Economy 122(6):1271–1319.
Galichon, A., and B. Salanié. 2022. Cupid's invisible hand: Social surplus and identification in matching models. Review of Economic Studies 89(5):2600–2629.
Goffman, E. 1963. Stigma: Notes on the Management of Spoiled Identity. Englewood Cliffs, NJ: Prentice-Hall.
Hausman, J. A., J. Abrevaya, and F. M. Scott-Morton. 1998. Misclassification of the dependent variable in a discrete-response setting. Journal of Econometrics 87(2):239–269.
Heckman, J., and R. Robb. 1985. Using longitudinal data to estimate age, period and cohort effects in earnings equations. In W. M. Mason and S. E. Fienberg, eds., Cohort Analysis in Social Research. New York: Springer.
Mahajan, A. 2006. Identification and estimation of regression models with misclassification. Econometrica 74(3):631–665.
Romano, J. P., and M. Wolf. 2005. Stepwise multiple testing as formalized data snooping. Econometrica 73(4):1237–1282.
Solon, G., S. J. Haider, and J. M. Wooldridge. 2015. What are we weighting for? Journal of Human Resources 50(2):301–316.
Tourangeau, R., and T. Yan. 2007. Sensitive questions in surveys. Psychological Bulletin 133(5):859–883.
Conclusion and Path Forward
The question at the centre of this paper is worth answering, and the strategy for answering it is original. Using a validated biological marker to study not just how much people misreport but what the misreporting reveals about the social costs they face is a genuine methodological contribution, and the demonstration that this matters for a substantive empirical question — who marries whom, and on what traits — is what distinguishes the paper from a purely technical exercise. The finding that measurement choice reshapes one side of the marriage market and not the other is striking and, if it survives scrutiny, portable well beyond smoking in Korea. Much of the data work is more careful than is common: the linkage rule is documented, the sample construction is laid out in a ladder, and losses to the linkage are diagnosed rather than ignored.
What is needed now is a revision that closes the distance between the measure and the interpretation. In order of priority, the author should establish what the concealment indicator captures by reporting the full cotinine distribution among concealers, re-estimating among wives of nonsmoking husbands, and reporting the gender gap under a smoker definition that drops the lifetime-cigarette screen that binds for women alone; should reconcile the sample definitions so that every reported figure traces to one consistently applied population, and should demonstrate that the three-group ordering and the husband-education gradient survive flexible age or cohort adjustment and an alternative parameterisation of the wife-education control; should re-run the Choo-Siow comparison on a common set of couples and develop the probability limits of the two matching estimators under an explicit misclassification process, which will likely support a stronger claim than the current one; should defend the time trend against the 2016 assay change, the 2010 questionnaire recoding, and the age-period-cohort identification problem; and should revisit the passages where underpowered or sign-reversed results are read as supportive. A compact framework for the disclosure decision would tie much of this together, converting asserted expectations into predictions capable of failing. These are demanding requests, but they are tractable with the data already in hand, and a paper that met them would make a contribution to both the measurement of social norms and the empirical literature on marriage matching.