Skip to content

LogicComm 0.12.0

Pre-release
Pre-release

Choose a tag to compare

@github-actions github-actions released this 09 Jun 08:01
b99a00d

LogicComm 0.12.0

Major change: cell-type co-expression scoring (neighborhood removal, stage 1)

LogicComm is pivoting to a pure cell-type-level cell-cell communication method,
positioned as an interpretable, sample-comparable alternative to CellChat. The
per-cell KNN/SNN neighborhood scoring is being removed because, for
dissociated scRNA-seq, that graph lives in expression space (transcriptomic
similarity), not physical space, and therefore cannot license spatial
juxtacrine/paracrine distance claims.

  • summarize_celltype_communication() now scores communication only at the
    cell-type level: for each sender -> receiver pair, the LCS is the fraction of
    the pair's opportunity universe (sender x receiver cell-count product) in which
    the ligand is active in the sender and the receptor is active in the receiver.
    The KNN/SNN neighborhood branch, edge construction, and graph helpers have been
    removed.
  • The legacy neighborhood arguments (knn_mat, graph_name, mode,
    remove_self_edges, graph_symmetrize, edge_weight_mode) are accepted via
    ... for backward compatibility but are deprecated and ignored with a
    warning
    . They will be removed in a later release.
  • The former neighborhood/local/distal/range output fields
    (lcs_neighborhood, local_active, distal_candidate, communication_range,
    ...) are retained as inert transitional stubs and will be removed once readers
    are migrated.

Stage 3: per-cell scoring engines de-neighborhooded

The standalone scoring engines turned out to be dual-mode (neighborhood +
global), and the flagship multi-sample pipeline run_multisample() is built on
IdentifyLogicConsensus(). Rather than hard-deleting and breaking the flagship,
their KNN/neighborhood paths were stripped and their global co-expression
cores retained
:

  • IdentifyLogicConsensus() is now global co-expression only (LCS = fraction of
    cells co-expressing the complete ligand and receptor logic). KNN graph
    resolution, symmetrization, edge construction, and weighted-edge scoring
    removed.
  • logic_score_lr() drops its mode/seurat_obj/knn_mat/graph arguments and
    scores globally (gate-aware path unchanged).
  • run_multisample() computes per-sample LCS by global co-expression; it no
    longer passes a per-sample KNN graph.
  • Legacy neighborhood arguments remain accepted via ... and ignored with a
    warning.

Stage 5: trustworthy statistics for selecting real communications

  • Axis-level permutation null. permute_celltype_communication() now
    defaults to an axis-level null (metric = "lcs"): every sender -> receiver ->
    L-R axis gets its own empirical p-value and BH FDR, instead of one
    cell-type-pair-level p-value being broadcast onto all of its L-R pairs. The
    output gains an lr_pair column; rank_communication_axes() joins the
    permutation evidence per axis and exposes permutation_fdr. Pair-level nulls
    (metric = "sum_lcs") remain available for backward compatibility. The
    cell-type co-expression pivot makes per-axis nulls cheap (no graph to rescan).
  • Per-cell bootstrap. bootstrap_celltype_communication() now resamples on
    the independent unit (cells) rather than the sender x receiver cell-count
    product. Co-expression LCS = (ligand-active fraction of sender cells) x
    (receptor-active fraction of receiver cells), so each fraction is bootstrapped
    as a binomial proportion over its own cell count. The previous approach used
    n_edges (the product) as the trial size and produced anticonservative,
    far-too-narrow intervals.

Standard discovery pipeline (footgun-free)

  • New discover_celltype_communication() wires the whole workflow in one call
    with the correct defaults: cell-type co-expression scoring, specificity +
    proliferation-confound annotation, an axis-level permutation null
    (metric = "lcs", so it is impossible to accidentally request a pair-level
    null via metric = "sum_lcs"), evidence ranking, a confound-filtered
    discovery view, and an FDR-passing shortlist. Pass expr = the counts matrix
    to enable the proliferation/breadth filter (cycling clusters otherwise dominate
    the ranking). A copy-paste template lives in
    inst/workflows/standard_discovery.R.

Stage 4: REO rank-weighted LCS (opt-in)

  • summarize_celltype_communication(lcs_weighting = "rank") scores by REO
    intensity instead of the binary co-expression fraction: the prevalence-weighted
    within-cell rank of the ligand in the sender type (fraction expressing times
    mean rank among expressers) times the same for the receptor. This restores
    dynamic range that binarizing discards (a ligand at the 99th within-cell
    percentile separates from one at the 51st) while keeping the prevalence signal
    that separates specific from ubiquitous axes, addressing the compressed,
    near-floor LCS values seen on real data. On the benchmark it edges out the
    binary score (AUPRC/sens@k); averaging rank among expressers only -- an earlier
    formulation the benchmark flagged -- discards prevalence and hurts precision.
    Requires the rank matrix from calc_REO_matrix(..., return_rank = TRUE). The
    default remains "binary", so existing results are unchanged.
  • The active call still uses the binary co-expression fraction, so the active
    axis set is identical under either weighting -- only the reported lcs differs.
  • permute_celltype_communication() inherits lcs_weighting from the scored
    object and carries the rank matrix through, so a rank-weighted observed score is
    tested against a rank-weighted null (coherent significance). It also flows
    through discover_celltype_communication(..., lcs_weighting = "rank") when the
    input is built with return_rank = TRUE.

Stage 6 (started): benchmark harness vs. baselines

  • inst/benchmark/benchmark_vs_baselines.R: a sandbox-runnable harness that
    simulates data with a known ground truth and the confounds that break naive
    scores -- ubiquitous/housekeeping pairs, a broad moderate-abundance pair, a
    CYCLING / transcriptional-breadth hub cell type that co-expresses many L-R
    genes, uneven cell-type sizes incl. a rare type, and true axes at three signal
    strengths. It scores every sender -> receiver -> L-R axis with LogicComm
    (binary, rank, and +proliferation-filter) and with re-implemented baselines
    (CellPhoneDB/CellChat-style mean-expression product; naive co-detection), and
    reports AUROC / AUPRC / sensitivity-at-k (random tie-breaking, averaged over
    sims). Real, runnable score_cellchat() and score_liana() adapters (and a
    note on CellPhoneDB via LIANA) are included for use where those packages exist.
  • Result (mean over sims; 1188 axes, 12 true positives): LogicComm with the
    proliferation filter reaches AUROC 0.995 / AUPRC 0.80, LogicComm (rank) 0.97 /
    0.63 and (binary) 0.99 / 0.58, while the mean-expression-product and naive
    baselines collapse to AUPRC ~0.04 -- ubiquitous, abundance and cycling-breadth
    confounds saturate their scores. The proliferation filter is the single biggest
    contributor to precision.

Still pending (non-blocking; the package is functional and green without them):
removal of the inert transitional stub columns (lcs_neighborhood,
communication_range, local_active, distal_candidate, ...) that the
cell-type path still emits as dead weight -- a cross-cutting refactor of the
downstream readers; and a possible re-architecture of IdentifyRankLogicConsensus()
to cell-type level (its graph-free mode is autocrine-only, so it keeps an optional
neighborhood mode for now). The gate-aware consensus and the spatial module retain
a cell-level graph by design. Also pending: real-package (CellChat/LIANA) and
real-data benchmark runs.