Package {transDIF}


Title: Score Comparability for Translated and Adapted Exams
Version: 0.1.0
Description: An adaptation-comparability workflow for small, lower-scoring and unbalanced language groups where standard differential item functioning (DIF) tools (Magis, Beland, Tuerlinckx and De Boeck, 2010, <doi:10.3758/BRM.42.3.847>) break down. Calibrates the Rasch model in each language group, links the groups robustly through the densest cluster of items rather than assuming DIF cancels out (on anchor selection see Kopf, Zeileis and Strobl, 2015, <doi:10.1177/0013164414529792>), detects small-sample DIF with an empirical-Bayes spike-and-slab model and local false discovery rates (Efron, 2004, <doi:10.1198/016214504000000089>), quantifies whether item-level DIF accumulates into different pass rates, explains DIF by item features to give translators actionable guidance, and drafts a comparability report.
License: MIT + file LICENSE
URL: https://github.com/edidatasolutions/transDIF, https://edidatasolutions.github.io/transDIF/
BugReports: https://github.com/edidatasolutions/transDIF/issues
Encoding: UTF-8
Depends: R (≥ 4.1)
Imports: stats
Suggests: knitr, markdown
VignetteBuilder: knitr
Config/roxygen2/version: 8.1.0
NeedsCompilation: no
Packaged: 2026-09-28 00:45:20 UTC; User
Author: Daniel Edi ORCID iD [aut, cre, cph]
Maintainer: Daniel Edi <danieledi2026@gmail.com>
Repository: CRAN
Date/Publication: 2026-10-08 09:20:02 UTC

Calibrate each language group separately

Description

Fits the Rasch model by marginal maximum likelihood within each group, each with its own mean-zero ability scale. The between-language difference 'd = b_focal - b_ref' therefore equals item DIF plus a common shift 'c = -(focal mean ability on the reference scale)'. Separating the two is the linking problem [td_dif()] solves.

Usage

td_calibrate(responses, group, ref = "ref", focal = "focal")

Arguments

responses

0/1 matrix (persons x items), column names = item ids.

group

Group label per person.

ref, focal

Labels of the reference (source-language) and focal (translated) groups.

Value

An 'td_calibration': data frame 'items' ('item', 'b_ref', 'se_ref', 'b_focal', 'se_focal', 'd', 'se_d') plus group sizes, ability SDs and ‘loc_var' (variance of the two scales’ locations). 'se_ref'/'se_focal' are absolute SEs; 'se_d' uses relative SEs, excluding the location uncertainty that is common to all items (it belongs to the linking shift).

Examples

sim <- td_simulate(n_ref = 400, n_focal = 120, n_items = 20, seed = 5)
cal <- td_calibrate(sim$responses, sim$group)
head(cal$items)

Robust linking, anchor selection and small-sample DIF

Description

The between-language difference of item 'i' is modeled as 'd_i ~ N(c + delta_i, se_i^2)', where 'c' is the common shift (the ability difference) that linking must recover and ‘delta_i' is the item’s DIF.

Usage

td_dif(
  calibration,
  fdr = 0.1,
  tau0 = 0.05,
  anchor_max = 0.2,
  link = c("mode", "mixture"),
  ambiguity = 0.8
)

Arguments

calibration

A 'td_calibration'.

fdr

Target Bayesian false discovery rate for flagging.

tau0

SD of DIF among "DIF-free" items: the scale of DIF considered negligible (logits).

anchor_max

Items with posterior DIF probability below this are reported as anchors.

link

'"mode"' (default) or '"mixture"'.

ambiguity

Competing-mode density ratio above which the linking is reported as ambiguous.

Details

**Linking.** By default 'c' is the precision-weighted *mode* of the 'd_i': the center of the densest cluster of items, found by maximizing 'sum_i phi((d_i - c) / s_i) / s_i' with 's_i^2 = se_i^2 + tau0^2'. It assumes that DIF-free items form the largest cluster, not that DIF cancels out as mean linking does. In known-truth benchmarks across balanced, directional and heavy DIF, it had the lowest overall RMSE of the estimators compared (including mean linking, iterative purification, median, Tukey biweight, least trimmed squares and the joint mixture below). Its SE is a sandwich estimate plus the scales' location variance. If a second cluster of items is nearly as dense, the linking is flagged as ambiguous and the competing shift is reported.

**DIF.** Given 'c', DIF effects follow a spike-and-slab mixture: DIF-free items (share 'pi0 > 0.5') have 'delta_i ~ N(0, tau0^2)', and DIF items have 'delta_i ~ N(m1, tau1^2)', where 'm1' allows directional DIF. Each item gets a posterior DIF probability, a shrunken DIF estimate and a local false discovery rate. 'link = "mixture"' instead estimates 'c' jointly in the mixture (the approach of version 0.1.0; kept for comparison).

Value

A 'td_dif' object: '$items' ('item', 'd', 'se_d', 'p_dif', 'lfdr', 'dif_mean', 'dif_sd' (posterior mean/SD of DIF), 'flag', 'anchor'), '$link' ('c', 'c_se', 'pi0', 'm1', 'tau0', 'tau1', 'alt_c', 'alt_ratio'), baseline linking constants 'c_mean' (all items as anchors) and 'c_purified', and 'weakly_identified' (TRUE, with a warning, when a second item cluster is nearly as dense as the chosen one).

Examples

sim <- td_simulate(n_ref = 400, n_focal = 120, n_items = 20, seed = 5)
dif <- td_dif(td_calibrate(sim$responses, sim$group))
dif
# true linking shift is -focal_mean = 0.5
c(estimate = dif$link[["c"]], mean_linking = dif$c_mean)

Which item features predict DIF?

Description

Random-effects meta-regression of each item's DIF estimate ('d - c') on item features (e.g. idioms, cultural referents, measurement units, vocabulary load), weighting by '1 / (se_d^2 + tau^2)'. The between-item variance 'tau^2' not explained by the features is estimated by the method of moments. Coefficients are logits of DIF per unit of the feature. That is guidance a translation team can act on.

Usage

td_features(dif, features)

Arguments

dif

An 'td_dif'.

features

Data frame with 'item' and numeric feature columns.

Value

Data frame of coefficients ('term', 'estimate', 'se', 'z', 'p_value') with attribute 'tau' (residual DIF SD). Features with no variation across items are dropped with a warning.

Examples

sim <- td_simulate(n_ref = 400, n_focal = 120, n_items = 20, seed = 5)
dif <- td_dif(td_calibrate(sim$responses, sim$group))
td_features(dif, sim$features)

Does item-level DIF add up to different pass rates?

Description

Differential test functioning for the translated form. With difficulties on the reference scale ('b_ref') and DIF effects 'delta', the translated form has difficulties 'b_ref + delta'. For the translated-language population (ability N(-c, sigma_focal^2) on the reference scale) the function computes the pass rate at raw cut 'cut' on the translated form versus on a DIF-free form, using exact score distributions. It also computes the expected raw-score shift for an examinee exactly at the cut. Intervals propagate both the linking uncertainty (the focal group's mean ability is '-c') and the posterior uncertainty of each item's DIF.

Usage

td_impact(dif, cut, n_draws = 200, level = 0.9, seed = NULL)

Arguments

dif

An 'td_dif'.

cut

Raw-score passing standard.

n_draws

Posterior draws.

level

Interval level.

seed

Optional seed.

Value

A data frame with estimate and interval for 'pass_rate_fair', 'pass_rate_translated', 'pass_rate_change' and 'score_shift_at_cut'.

Examples

sim <- td_simulate(n_ref = 400, n_focal = 120, n_items = 20, seed = 5)
dif <- td_dif(td_calibrate(sim$responses, sim$group))
td_impact(dif, cut = 12, n_draws = 50, seed = 1)

Mantel-Haenszel DIF with purification (conventional baseline)

Description

Mantel-Haenszel DIF with purification (conventional baseline)

Usage

td_mh(responses, group, ref = "ref", focal = "focal", fdr = 0.1, purify = TRUE)

Arguments

responses

0/1 matrix.

group

Group label per person.

ref, focal

Group labels.

fdr

Benjamini-Hochberg level for flagging.

purify

Re-run once matching on the total over non-flagged items.

Value

Data frame: 'item', 'alpha_mh', 'delta_mh' (ETS delta scale), 'p_value', 'flag'.

Examples

sim <- td_simulate(n_ref = 400, n_focal = 120, n_items = 20, seed = 5)
mh <- td_mh(sim$responses, sim$group)
table(flagged = mh$flag, true_dif = sim$truth$dif_item)

Draft a comparability report

Description

Assembles a plain-language Markdown report of the linking, DIF, aggregate impact and feature analyses, organized around the kinds of evidence the Standards for Educational and Psychological Testing (2014; fairness chapter) and the ITC Guidelines for Translating and Adapting Tests (2nd ed., 2017) ask for. It is a draft for a psychometrician to review, not a finished validity argument.

Usage

td_report(
  dif,
  impact = NULL,
  features = NULL,
  languages = c("source", "target"),
  file = NULL
)

Arguments

dif

An 'td_dif'.

impact

Optional output of [td_impact()].

features

Optional output of [td_features()].

languages

Names of the source and target languages.

file

Optional path to write the Markdown to.

Value

The report as a character string (invisibly if written to file).

Examples

sim <- td_simulate(n_ref = 400, n_focal = 120, n_items = 20, seed = 5)
dif <- td_dif(td_calibrate(sim$responses, sim$group))
rep <- td_report(dif, td_impact(dif, cut = 12, n_draws = 50, seed = 1),
                 td_features(dif, sim$features), c("English", "French"))
cat(rep)
# To save: td_report(dif, file = file.path(tempdir(), "comparability.md"))

Simulate a source-language and a translated administration with known DIF

Description

Items carry binary adaptation features (idiom, cultural referent, measurement units, high vocabulary load). In the translated form an item's difficulty shifts by 'sum(feature_effects * features)' plus small noise, so DIF is directional (translations mostly harder) and unbalanced, which is the case where mean-based linking fails. The focal (translated) group is small and lower-scoring on average.

Usage

td_simulate(
  n_ref = 2000,
  n_focal = 150,
  n_items = 60,
  focal_mean = -0.5,
  focal_sd = 1,
  feature_prev = c(idiom = 0.1, cultural = 0.1, units = 0.08, vocabulary = 0.12),
  feature_effects = c(idiom = 0.6, cultural = 0.5, units = -0.4, vocabulary = 0.35),
  dif_noise = 0.1,
  seed = NULL
)

Arguments

n_ref, n_focal

Group sizes.

n_items

Test length.

focal_mean, focal_sd

Focal-group ability (reference is N(0, 1)).

feature_prev

Prevalence of each feature.

feature_effects

DIF (logits) contributed by each feature.

dif_noise

SD of feature-unrelated DIF on flagged items.

seed

Optional seed.

Value

An 'td_sim': '$responses' (0/1 matrix, persons x items), '$group' ('"ref"'/'"focal"'), '$features' (item data frame), '$truth' ('b_ref', 'dif', 'focal_mean', 'focal_sd').

Examples

sim <- td_simulate(n_ref = 400, n_focal = 120, n_items = 20, seed = 5)
table(dif_item = sim$truth$dif_item)
head(sim$features)