API Reference
Factorization
scrise.factorization
pf2
pf2(
X: AnnData | None = None,
rank: int | None = None,
random_state=1,
doEmbedding: bool = True,
tolerance=1e-06,
max_iter: int = 100,
normalize_slices: bool = False,
backend: str | None = None,
compress: int
| tuple[int, int | None]
| str
| bool
| None = None,
compression_kwarg: dict[str, Any] | None = None,
parafac2_kwarg: dict[str, Any] | None = None,
condition_key: str | None = None,
adata: AnnData | None = None,
) -> anndata.AnnData
Perform PARAFAC2 tensor decomposition on single-cell RNA-seq data.
This is the main function for running RISE analysis. It decomposes the multi-condition single-cell data into condition factors, eigen-state factors, and gene factors, revealing patterns across experimental conditions.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
AnnData
|
Preprocessed AnnData object containing single-cell RNA-seq data. Must have X.obs["condition_unique_idxs"] indicating which condition each cell belongs to (0-indexed). |
None
|
rank
|
int
|
Number of components to extract. Determines the complexity of the decomposition. Typically chosen based on variance explained and Factor Match Score analysis (see plot_r2x and plot_fms_diff_ranks). |
None
|
random_state
|
int, optional (default: 1)
|
Random seed for reproducibility of the decomposition. |
1
|
doEmbedding
|
bool, optional (default: True)
|
If True, automatically computes PaCMAP embedding of cell projections and stores in X.obsm["X_pf2_PaCMAP"]. This enables visualization functions like plot_labels_pacmap. |
True
|
tolerance
|
float, optional (default: 1e-6)
|
Convergence threshold for the optimization algorithm. Lower values increase precision but may require more iterations. |
1e-06
|
max_iter
|
int, optional (default: 100)
|
Maximum number of iterations for the optimization algorithm. |
100
|
normalize_slices
|
bool, optional (default: False)
|
If True, weights each condition by the inverse of its Frobenius norm, so conditions count equally regardless of cell count. Factors and R2X stay in data units, comparable to an unweighted fit. |
False
|
backend
|
str | None, optional (default: None)
|
Compute backend to run matrix products on: one of |
None
|
compress
|
int | tuple[int, int | None] | str | bool | None, optional (default: None)
|
CANDELINC compression mode passed to |
None
|
compression_kwarg
|
dict
|
Additional keyword arguments forwarded to
:func: |
None
|
parafac2_kwarg
|
dict
|
Additional keyword arguments forwarded to |
None
|
condition_key
|
str, optional (default: None)
|
Column in |
None
|
adata
|
anndata.AnnData, optional (default: None)
|
Alias for |
None
|
Returns:
| Type | Description |
|---|---|
AnnData
|
The input AnnData object with added RISE decomposition results:
Components are ordered from highest to lowest intrinsic energy
(see :func: |
Source code in scrise/factorization.py
223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 | |
correct_conditions
Correct the condition factors by normalizing for overall read depth.
This function adjusts condition factors (stored in X.uns["Pf2_A"]) to account for differences in sequencing depth across conditions. It uses linear regression to model the relationship between total read counts and condition factor magnitudes, then applies a correction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
AnnData
|
AnnData object containing RISE decomposition results. Must have: - X.obs["condition_unique_idxs"]: 0-indexed condition assignments - X.uns["Pf2_A"]: Condition factors from PARAFAC2 decomposition |
required |
Returns:
| Type | Description |
|---|---|
ndarray
|
Corrected condition factors normalized by sequencing depth |
Source code in scrise/factorization.py
order_components_by_energy
Reorder PARAFAC2 components by intrinsic energy.
RISE previously inherited an ordering of components based on the Gini coefficient (variance-to-mean ratio) of the condition factor. That ordering is a heuristic: it is not particularly stable, and it does not guarantee that, when moving from rank N to rank N+1, the first N components of the new fit correspond to the N components of the old fit.
This function instead orders components by their intrinsic energy,
|weights[r]| * ||A[:, r]|| * ||C[:, r]|| (the product of component
weights, condition-factor, and gene-factor column norms), which is
directly determined by the fit and is not subject to arbitrary rescaling.
Components are ordered from highest to lowest energy, so that low-energy
components -- which tend to be the ones added when the rank is
increased -- land at the high end of the ordering. This is consistent
with the expectation, motivated by Harshman's Uniqueness Theorem for
PARAFAC2 (which fixes the decomposition up to permutation and
sign/scale given >= 3 conditions and full-column-rank factors), that a
rank-(N+1) fit should mostly agree with a rank-N fit on its first N
components.
Before computing the ordering, a canonical sign is imposed on each
component (see :func:canonical_component_signs) so that the ordering,
and any downstream comparison across fits, is well defined.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
AnnData
|
AnnData object containing RISE decomposition results (as produced by
:func: |
required |
Returns:
| Type | Description |
|---|---|
AnnData
|
The same AnnData object, with Pf2_A, Pf2_B, Pf2_C, Pf2_weights
(and, if present, obsm["projections"] and
obsm["weighted_projections"]) reordered/updated in place. The
eigen-state axis of Pf2_B, and the matching columns of
obsm["projections"], are permuted alongside the components so that
the maximal-diagonal form of B established by
|
Source code in scrise/factorization.py
87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 | |
canonical_component_signs
Compute a canonical sign for each component of a gene factor matrix.
PARAFAC2 components are only identified up to a sign flip that is shared
between two of the three factor matrices (Harshman's Uniqueness Theorem
fixes the decomposition up to permutation and sign/scale). This makes any
comparison across fits (different ranks, different random seeds, etc.)
ambiguous unless a canonical sign is first imposed. We adopt the
convention that the largest-magnitude entry of each gene-factor (C)
column should be positive.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
C
|
ndarray
|
Gene factor matrix, shape (n_genes, rank). |
required |
Returns:
| Type | Description |
|---|---|
ndarray
|
Array of +1/-1 signs, shape (rank,), one per component. |
Source code in scrise/factorization.py
match_components_across_ranks
match_components_across_ranks(
C_low: ndarray, C_high: ndarray, threshold: float = 0.6
) -> tuple[np.ndarray, np.ndarray]
Match components between two PARAFAC2 fits of adjacent rank by cosine similarity of their gene factors.
This provides the "does component X at rank N correspond to component Y
at rank N-1" matching primitive described in issue #520. It is a small,
self-contained addition, not a full cross-rank benchmarking pipeline:
given the gene factors of a rank-N fit and a rank-(N+1) fit (each
already sign-canonicalized, e.g. via :func:canonical_component_signs),
it performs Hungarian maximum-weight matching on cosine similarity and
reports which components matched, and which rank-(N+1) component(s) had
no good match (i.e. are candidates for the newly added component).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
C_low
|
ndarray
|
Gene factor matrix of the lower-rank fit, shape (n_genes, rank_low). |
required |
C_high
|
ndarray
|
Gene factor matrix of the higher-rank fit, shape (n_genes, rank_high), with rank_high >= rank_low. |
required |
threshold
|
float, optional (default: 0.6)
|
Minimum cosine similarity for two components to be considered a match. |
0.6
|
Returns:
| Type | Description |
|---|---|
tuple of numpy.ndarray
|
matched_pairs : numpy.ndarray of shape (n_matches, 2)
Each row is (index in C_low, index in C_high) for a matched
pair of components.
unmatched_high : numpy.ndarray
Indices into C_high of components with no match above
|
Source code in scrise/factorization.py
Factor Import/Export
scrise.factor_io
Reading and writing RISE factors, without the raw expression matrix.
Factors are exported to h5ad with the projection matrix OPQ-quantized and the cell barcodes packed into a uint8 matrix, both of which have to be undone on load.
export_factors
export_factors(
X: AnnData,
filename: str,
fidelity_threshold: float = 0.99,
random_state: int = 42,
) -> anndata.AnnData
Export RISE decomposition factors to an h5ad file without raw expression data.
Compresses the projection matrix using Optimized Product Quantization (OPQ) to meet or exceed the specified fidelity threshold (R^2 >= fidelity_threshold). All factor matrices (Pf2_A, Pf2_B, Pf2_weights, Pf2_C) are stored in float32. Weighted projections are never stored on disk because they can be reconstructed deterministically as projections @ Pf2_B. PaCMAP embeddings are optionally stored if present.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
AnnData
|
AnnData object containing RISE decomposition results. Must contain: - X.uns["Pf2_A"], X.uns["Pf2_B"], X.uns["Pf2_weights"] - X.varm["Pf2_C"] - X.obsm["projections"] |
required |
filename
|
str
|
Output file path (.h5ad). |
required |
fidelity_threshold
|
float, optional (default: 0.99)
|
Target R^2 reconstruction accuracy threshold for projection compression. |
0.99
|
random_state
|
int, optional (default: 42)
|
Random seed for reproducibility during OPQ codebook training. |
42
|
Returns:
| Type | Description |
|---|---|
AnnData
|
The factor-only AnnData object written to disk. |
Source code in scrise/factor_io.py
205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 | |
load_factors
Load RISE decomposition factors from an h5ad file, decompressing OPQ projections and optionally rebuilding the full dataset from raw expression data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
filename
|
str
|
Path to factors .h5ad file. |
required |
raw_path
|
str
|
Path to raw AnnData or IVCSR .h5ad / .h5 file. If provided, cells and genes are matched based on cell barcodes and gene names to attach the expression matrix. |
None
|
Returns:
| Type | Description |
|---|---|
AnnData
|
AnnData object with reconstructed projections, weighted_projections, factors, and optionally the raw expression matrix X. |
Source code in scrise/factor_io.py
Rank Selection
scrise.rank_selection
Rank selection for RISE via bi-cross-validation (BiCV).
Bi-cross-validation extends ordinary cross-validation to two-way (row and column) held-out blocks. For RISE, we hold out a random subset of cells and a random subset of genes, fit PARAFAC2 on the remaining (train-cell x train-gene) block, and then measure how well the fitted model predicts the held-out (test-cell x test-gene) block. Unlike the ordinary in-sample fit R2X (which increases monotonically with rank), the BiCV R2X penalizes overfitting and typically peaks near the "true" rank of the data.
bicv
bicv(
X: AnnData | None = None,
ranks: Sequence[int] | None = None,
n_repeats: int = 3,
held_out_cell_frac: float = 0.5,
held_out_gene_frac: float = 0.5,
random_state: int | None = None,
tolerance: float = 1e-06,
max_iter: int = 200,
compress: int
| tuple[int, int | None]
| str
| bool
| None = "auto",
compression_kwarg: dict[str, Any] | None = None,
parafac2_kwarg: dict[str, Any] | None = None,
condition_key: str | None = None,
adata: AnnData | None = None,
) -> pd.DataFrame
Evaluate rank via bi-cross-validation (BiCV) and in-sample fit R2X.
For each candidate rank, computes both the ordinary in-sample fit R2X
(using the full dataset, as in :func:scrise.factorization.rise_pca_r2x)
and the BiCV R2X (n_repeats independent random cell/gene splits,
each evaluated at every rank -- see :func:_bicv_trial). The fit R2X
increases monotonically with rank; the BiCV R2X penalizes overfitting and
typically peaks near the rank that best generalizes to held-out data.
Plot both with :func:scrise.plotting.plot_bicv_r2x to select a rank.
Each of the n_repeats splits (and the full dataset, for the in-sample
fit) is compressed once -- sized for the largest rank requested -- and
every rank's PARAFAC2 fit reuses that same compression, rather than
recompressing per rank (see :func:_fit_at_ranks). Compression is by far
the most expensive, O(nnz) step against the raw data; fitting a
smaller rank from an already-compressed representation only touches the
small dense compressed cores. This applies whenever compression_kwarg
is given (as it must be to control e.g. n_power_iter); without it,
each rank still goes through its own call into parafac2_nd's internal
compression shortcut, unchanged.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
AnnData
|
Preprocessed AnnData object containing single-cell RNA-seq data.
Must have X.obs["condition_unique_idxs"] and X.var["means"]
(as produced by |
None
|
ranks
|
sequence of int
|
Candidate rank values to evaluate (e.g., [5, 10, 15, 20, 25, 30]). |
None
|
n_repeats
|
int, optional (default: 3)
|
Number of independent random cell/gene splits, each evaluated at
every rank in |
3
|
held_out_cell_frac
|
float, optional (default: 0.5)
|
Fraction of cells held out per condition in each BiCV trial. |
0.5
|
held_out_gene_frac
|
float, optional (default: 0.5)
|
Fraction of genes held out in each BiCV trial. |
0.5
|
random_state
|
int
|
Random seed for reproducibility. |
None
|
tolerance
|
float, optional (default: 1e-6)
|
Convergence threshold passed to the PARAFAC2 fit. |
1e-06
|
max_iter
|
int, optional (default: 200)
|
Maximum number of iterations passed to the PARAFAC2 fit. |
200
|
compress
|
int | tuple[int, int | None] | str | bool | None
|
CANDELINC compression mode passed to each PARAFAC2 fit. Defaults to
|
'auto'
|
compression_kwarg
|
dict
|
Additional keyword arguments forwarded to
:func: |
None
|
parafac2_kwarg
|
dict
|
Additional keyword arguments forwarded to |
None
|
condition_key
|
str, optional (default: None)
|
Column in |
None
|
adata
|
anndata.AnnData, optional (default: None)
|
Alias for |
None
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
Long-form DataFrame with columns "Rank", "Repeat", "Metric" (one of
"Fit R2X" or "BiCV R2X"), and "R2X". Ready to pass to
:func: BiCV rows carry per-trial diagnostics as additional columns:
These columns are NaN on "Fit R2X" rows, which come from an unsplit fit on the full dataset. |
Source code in scrise/rank_selection.py
504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 | |
Annotation Alignment
scrise.annotation_alignment
Cell-type alignment scoring for RISE component projections.
Quantifies how well a RISE component's cell loadings / projections align with annotated cell types, answering: 1. Uniqueness: Does the component concentrate on a single annotated cell type (tau)? 2. Combination alignment: Does it align with a specific subset of cell types (AUROC + FDR, eta^2)?
CellTypeAlignmentResults
dataclass
Alignment results across multiple RISE components.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
results
|
list[ComponentAlignmentResult]
|
List of alignment results for individual components. |
required |
enrichment
|
DataFrame
|
Components x cell types matrix of AUROC values. |
required |
p_values
|
DataFrame
|
Components x cell types matrix of permutation p-values. |
required |
q_values
|
DataFrame
|
Components x cell types matrix of joint BH FDR-adjusted p-values. |
required |
tau
|
Series
|
Tissue specificity index tau for each component. |
required |
eta_squared
|
Series
|
Eta-squared (variance explained) for each component. |
required |
kruskal_epsilon_squared
|
Series
|
Kruskal-Wallis epsilon-squared for each component. |
required |
significant_cell_types
|
dict[int | str, list[str]]
|
Mapping of component to significant cell types. |
required |
alpha
|
float
|
Significance threshold used for q-values. |
0.05
|
Source code in scrise/annotation_alignment.py
summary
Return a summary DataFrame across all components.
Source code in scrise/annotation_alignment.py
ComponentAlignmentResult
dataclass
Alignment results for a single RISE component.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
component
|
int | str
|
Component index or label. |
required |
enrichment
|
Series
|
AUROC per cell type distinguishing that type from all others. |
required |
p_values
|
Series
|
Empirical permutation p-values for enrichment (AUROC > null). |
required |
q_values
|
Series
|
Benjamini-Hochberg FDR-adjusted p-values. |
required |
tau
|
float
|
Tissue specificity index tau (uniqueness score in [0, 1]). |
required |
eta_squared
|
float
|
Proportion of loading variance explained by cell type (in [0, 1]). |
required |
kruskal_epsilon_squared
|
float
|
Non-parametric effect size epsilon-squared from Kruskal-Wallis. |
required |
significant_cell_types
|
list[str]
|
List of cell types with significant enrichment (q <= alpha and AUROC > 0.5). |
list()
|
alpha
|
float
|
Significance threshold used for q-values. |
0.05
|
Source code in scrise/annotation_alignment.py
to_dict
Convert result to a dictionary.
Source code in scrise/annotation_alignment.py
cell_type_alignment
cell_type_alignment(
loadings: ndarray | Series,
cell_types: Series | ndarray,
signed: bool = False,
n_permutations: int = 1000,
alpha: float = 0.05,
random_state: int | Generator | None = None,
component_label: int | str = 1,
) -> ComponentAlignmentResult
Score alignment of a single component's cell loadings with annotated cell types.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
loadings
|
ndarray | Series
|
Cell loading vector for one component (shape: n_cells,). |
required |
cell_types
|
Series | ndarray
|
Cell type annotations (length: n_cells). |
required |
signed
|
bool, optional (default: False)
|
If True, uses the absolute value of loadings (|loading|), which is appropriate for signed eigen-state projections (e.g. P_i @ B[:, r]). |
False
|
n_permutations
|
int, optional (default: 1000)
|
Number of label permutations to compute empirical p-values. If 0, p-values are computed using the asymptotic one-sided Mann-Whitney U test. |
1000
|
alpha
|
float, optional (default: 0.05)
|
FDR significance threshold for identifying enriched cell types. |
0.05
|
random_state
|
int | np.random.Generator | None, optional (default: None)
|
Random seed or Generator for permutation reproducibility. |
None
|
component_label
|
int | str, optional (default: 1)
|
Identifier for the component. |
1
|
Returns:
| Type | Description |
|---|---|
ComponentAlignmentResult
|
Dataclass containing AUROC enrichment, p-values, q-values, tau, eta^2, and significant cell types. |
Source code in scrise/annotation_alignment.py
201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 | |
score_cell_type_alignment
score_cell_type_alignment(
data: AnnData | ndarray | DataFrame,
cell_types: Series | ndarray | str | None = None,
signed: bool = False,
projection_key: str = "weighted_projections",
n_permutations: int = 1000,
alpha: float = 0.05,
random_state: int | Generator | None = None,
) -> CellTypeAlignmentResults
Score cell-type alignment across all RISE components.
Jointly calculates per-cell-type AUROC enrichment, empirical significance with Benjamini-Hochberg FDR correction across all (component x cell type) tests, uniqueness (tau), and combination alignment (eta^2).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
AnnData | ndarray | DataFrame
|
AnnData containing fitted RISE results, or a matrix of cell loadings (shape: n_cells, n_components). |
required |
cell_types
|
Series | ndarray | str | None
|
Cell type annotations. If data is AnnData and cell_types is a string (or None), looks up data.obs[cell_types] (defaults to 'cell_type' or 'CellType'). |
None
|
signed
|
bool, optional (default: False)
|
If True, takes the absolute value of loadings (|loading|). |
False
|
projection_key
|
str, optional (default: "weighted_projections")
|
Key in data.obsm to extract loadings from when data is an AnnData. Defaults to 'weighted_projections', or falls back to 'projections'. |
'weighted_projections'
|
n_permutations
|
int, optional (default: 1000)
|
Number of permutations for null AUROC distribution. |
1000
|
alpha
|
float, optional (default: 0.05)
|
Significance threshold for FDR q-values. |
0.05
|
random_state
|
int | np.random.Generator | None, optional (default: None)
|
Random seed or Generator for permutations. |
None
|
Returns:
| Type | Description |
|---|---|
CellTypeAlignmentResults
|
Container with full results across all components. |
Source code in scrise/annotation_alignment.py
331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 | |
Alignment Statistics
scrise.alignment_stats
compute_auroc_per_cell_type
compute_auroc_per_cell_type(
loadings: ndarray,
cell_type_codes: ndarray,
n_types: int,
) -> np.ndarray
Compute AUROC for each cell type vs all other cells.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
loadings
|
ndarray
|
1D array of cell loadings of shape (n_cells,). |
required |
cell_type_codes
|
ndarray
|
1D array of integer cell type assignments in [0, n_types - 1]. |
required |
n_types
|
int
|
Total number of unique cell types. |
required |
Returns:
| Type | Description |
|---|---|
ndarray
|
1D array of AUROC values for each cell type of shape (n_types,). |
Source code in scrise/alignment_stats.py
compute_tau
Compute tissue specificity index tau (Yanai et al., 2005).
tau = sum(1 - x_hat) / (n_types - 1), where x_hat = x / max(x).
When computed over AUROC enrichment values: - tau -> 1 indicates the component is specific/private to a single cell type. - tau -> 0 indicates the component is evenly distributed across cell types.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
enrichment_scores
|
ndarray | Series
|
1D array of enrichment values (e.g. AUROC) across cell types. |
required |
baseline
|
float, optional (default: 0.0)
|
Baseline value subtracted before calculating tau. Values below baseline are clipped to 0. |
0.0
|
Returns:
| Type | Description |
|---|---|
float
|
Tau index in [0.0, 1.0]. |
Source code in scrise/alignment_stats.py
compute_eta_squared
Compute omnibus eta-squared (loading ~ cell_type).
Proportion of loading variance explained by cell-type identity.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
loadings
|
ndarray
|
1D array of cell loadings (shape: n_cells,). |
required |
cell_type_codes
|
ndarray
|
1D array of integer cell type assignments (shape: n_cells,). |
required |
n_types
|
int
|
Number of unique cell types. |
required |
Returns:
| Type | Description |
|---|---|
float
|
Eta-squared value in [0.0, 1.0]. |
Source code in scrise/alignment_stats.py
compute_kruskal_epsilon_squared
compute_kruskal_epsilon_squared(
loadings: ndarray,
cell_type_codes: ndarray,
n_types: int,
) -> float
Compute Kruskal-Wallis epsilon-squared effect size.
Non-parametric measure of association between loading and cell type.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
loadings
|
ndarray
|
1D array of cell loadings (shape: n_cells,). |
required |
cell_type_codes
|
ndarray
|
1D array of integer cell type assignments (shape: n_cells,). |
required |
n_types
|
int
|
Number of unique cell types. |
required |
Returns:
| Type | Description |
|---|---|
float
|
Epsilon-squared value in [0.0, 1.0]. |
Source code in scrise/alignment_stats.py
Quantization & Compression
scrise.opq
Optimized Product Quantization (OPQ) for projecting and compressing factor matrices.
OPQQuantizer
Optimized Product Quantizer for compressing projection matrices.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
M
|
int
|
Number of sub-quantizers (sub-vector partitions). |
required |
n_bits
|
int, optional (default: 8)
|
Number of bits per sub-quantizer. 8 bits corresponds to 256 centroids per sub-space. |
8
|
n_iter
|
int, optional (default: 5)
|
Number of alternating OPQ optimization iterations for the rotation matrix. |
5
|
random_state
|
int, optional (default: 42)
|
Random seed for reproducibility. |
42
|
Source code in scrise/opq.py
10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 | |
decode
Decode uint8 sub-quantizer codes back to the continuous feature space.
Source code in scrise/opq.py
encode
Encode data matrix X into uint8 sub-quantizer centroid indices.
Source code in scrise/opq.py
fit
Fit rotation matrix and sub-quantizer codebooks on data matrix X.
Source code in scrise/opq.py
fit_transform
Fit OPQ, encode X to codes, reconstruct, and compute R^2.
Source code in scrise/opq.py
from_saved
classmethod
from_saved(
R: ndarray,
centroids_cat: ndarray,
sub_dims: ndarray,
n_bits: int = 8,
) -> OPQQuantizer
Instantiate a fitted OPQQuantizer from saved parameters.
Source code in scrise/opq.py
find_optimal_opq
find_optimal_opq(
P: ndarray,
fidelity_threshold: float = 0.99,
random_state: int = 42,
) -> tuple[OPQQuantizer, np.ndarray, float]
Find the smallest number of sub-quantizers M achieving R^2 >= fidelity_threshold.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
P
|
ndarray
|
Projection matrix of shape (N, D). |
required |
fidelity_threshold
|
float, optional (default: 0.99)
|
Target R^2 reconstruction accuracy. |
0.99
|
random_state
|
int, optional (default: 42)
|
Random seed. |
42
|
Returns:
| Type | Description |
|---|---|
tuple of (OPQQuantizer, np.ndarray, float)
|
(quantizer, codes, r2) |
Source code in scrise/opq.py
Preprocessing
parafac2.normalize
Dataset preprocessing and normalization utilities for PARAFAC2 analysis.
This module provides functions to filter, normalize, and annotate single-cell gene expression datasets stored in AnnData objects prior to PARAFAC2 matrix factorization.
prepare_dataset
Preprocess and normalize an AnnData dataset for PARAFAC2 factorization.
Performs quality control filtering of low-count cells and low-expression genes, normalizes total cell counts and gene sums, applies a log10 transformation, and computes metadata required by PARAFAC2 (condition indices and gene means).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
AnnData
|
Input single-cell dataset with raw count matrix stored in |
required |
condition_name
|
str
|
Column name in |
required |
geneThreshold
|
float
|
Minimum threshold fraction for gene inclusion. Genes with total counts
less than |
required |
Returns:
| Type | Description |
|---|---|
AnnData
|
A filtered and normalized copy of the AnnData object. Contains the
log-transformed normalized counts in |
Source code in parafac2/normalize.py
Visualization Functions
General Plotting
scrise.plotting.general
plot_r2x
Plot variance explained (R²X) for RISE and PCA across different ranks.
This visualization helps determine the optimal number of components by showing how variance explained increases with rank. The elbow point where the curve flattens indicates a good balance between model complexity and explanatory power.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
AnnData
|
Preprocessed AnnData object containing single-cell RNA-seq data. Must have X.obs["condition_unique_idxs"] for RISE decomposition. |
required |
rank_vec
|
array-like of int
|
Array of rank values to test (e.g., [1, 5, 10, 15, 20, 25, 30]). Each rank represents a different number of components. |
required |
ax
|
Axes
|
Matplotlib axes object to plot on. |
required |
compress
|
int | tuple[int, int | None] | str | bool | None
|
CANDELINC compression mode passed to |
'auto'
|
Source code in scrise/plotting/general.py
Factor Plotting
scrise.plotting.factors
plot_condition_factors
plot_condition_factors(
data: AnnData,
ax: Axes,
cond: str = "Condition",
log_transform: bool = True,
cond_group_labels: Series | None = None,
ThomsonNorm: bool = False,
color_key=None,
group_cond: bool = False,
control_pattern: str | None = None,
control_conditions: Sequence[str] | None = None,
)
Plot condition factors as a heatmap showing how conditions contribute to components.
This visualization shows how each experimental condition (rows) contributes to each RISE component (columns). High values indicate strong association between a condition and a component's pattern. Log transformation and normalization help reveal relative differences across conditions.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
AnnData
|
AnnData object with RISE decomposition results. Must contain: - data.uns["Pf2_A"]: Condition factors (n_conditions, rank) - data.obs[cond]: Condition labels for each cell |
required |
ax
|
Axes
|
Matplotlib axes object to plot on. |
required |
cond
|
str, optional (default: "Condition")
|
Name of column in data.obs containing condition labels. |
'Condition'
|
log_transform
|
bool, optional (default: True)
|
If True, applies log10 transformation to condition factors before plotting. This helps visualize differences when values span orders of magnitude. |
True
|
cond_group_labels
|
pandas.Series, optional (default: None)
|
Series mapping conditions to group labels for colored row annotations. Useful for grouping related conditions (e.g., drug classes, patient cohorts). |
None
|
ThomsonNorm
|
bool, optional (default: False)
|
If True, normalizes factors using only control conditions (those containing 'CTRL'). |
False
|
color_key
|
list, optional (default: None)
|
Custom colors for condition group labels. If None, uses default palette. |
None
|
group_cond
|
bool, optional (default: False)
|
If True and cond_group_labels provided, sorts conditions by group. |
False
|
control_pattern
|
str, optional (default: None)
|
Substring identifying the control conditions whose median and spread
set the normalization. Overrides |
None
|
control_conditions
|
Sequence[str], optional (default: None)
|
Explicit control condition names, taking precedence over
|
None
|
Source code in scrise/plotting/factors.py
71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 | |
plot_eigenstate_factors
Plot eigen-state factors as a heatmap showing cell state patterns.
Eigen-state factors represent the underlying cell state patterns across components. Each row represents an eigen-state (a summary of similar cells), and each column represents a component. High values indicate strong association between a cell state pattern and a component.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
AnnData
|
AnnData object with RISE decomposition results. Must contain: - data.uns["Pf2_B"]: Eigen-state factors (rank, rank) |
required |
ax
|
Axes
|
Matplotlib axes object to plot on. |
required |
Source code in scrise/plotting/factors.py
plot_gene_factors
Plot gene factors as a heatmap showing which genes contribute to each component.
This visualization reveals coordinated gene modules by showing which genes (rows) are highly weighted in each component (columns). The weight parameter filters out genes with low contributions, focusing on the most important genes for interpretation.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
AnnData
|
AnnData object with RISE decomposition results. Must contain: - data.varm["Pf2_C"]: Gene factors (n_genes, rank) |
required |
ax
|
Axes
|
Matplotlib axes object to plot on. |
required |
weight
|
float, optional (default: 0.08)
|
Minimum absolute weight threshold for including genes. Genes with maximum absolute weight below this value across all components are filtered out. Higher values show fewer, more important genes. |
0.08
|
trim
|
bool, optional (default: True)
|
If True, filters genes based on the weight parameter. If False, shows all genes. |
True
|
Source code in scrise/plotting/factors.py
PaCMAP Visualization
scrise.plotting.pacmap
plot_labels_pacmap
plot_labels_pacmap(
X: AnnData,
labelType: str,
ax: Axes,
condition=None,
cmap: str = "tab20",
color_key=None,
)
Plot PaCMAP embedding colored by categorical labels (cell type or condition).
This visualization shows the overall structure of the cell embedding, revealing how cells cluster by cell type, experimental condition, or other categorical metadata. Useful for understanding the biological organization captured by RISE.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
AnnData
|
AnnData object with RISE decomposition results. Must contain: - X.obsm["X_pf2_PaCMAP"]: PaCMAP embedding coordinates (n_cells, 2) - X.obs[labelType]: Categorical labels for coloring cells |
required |
labelType
|
str
|
Name of column in X.obs containing categorical labels to color by. Common values: "Cell Type", "Condition", "Sample", etc. |
required |
ax
|
Axes
|
Matplotlib axes object to plot on. |
required |
condition
|
list of str, optional (default: None)
|
If provided, only highlights cells from these specific conditions/labels. All other cells are labeled as "Other". |
None
|
cmap
|
str, optional (default: "tab20")
|
Matplotlib colormap name for coloring categories. |
'tab20'
|
color_key
|
list, optional (default: None)
|
Custom list of colors for categories. If None, uses cmap. |
None
|
Source code in scrise/plotting/pacmap.py
plot_gene_pacmap
Plot PaCMAP embedding colored by gene expression levels.
This visualization overlays gene expression onto the PaCMAP embedding of cells, revealing which cell populations express specific genes. Useful for validating component interpretations by checking if marker genes align with component patterns.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
gene
|
str
|
Name of gene to visualize. Must be present in X.var_names. |
required |
X
|
AnnData
|
AnnData object with RISE decomposition results. Must contain: - X.obsm["X_pf2_PaCMAP"]: PaCMAP embedding coordinates (n_cells, 2) - X[:, gene]: Gene expression values - X.var["means"]: Pre-computed gene means for centering |
required |
ax
|
Axes
|
Matplotlib axes object to plot on. |
required |
clip_outliers
|
float, optional (default: 0.9995)
|
Quantile threshold for clipping extreme expression values. Values above this quantile are clipped to improve visualization contrast. |
0.9995
|
Source code in scrise/plotting/pacmap.py
plot_wp_pacmap
Plot PaCMAP embedding colored by weighted projections for a component.
This visualization shows which cells contribute most strongly to a specific component by coloring them according to their weighted projections. Cells with high weighted projections (bright colors) are most representative of that component's expression pattern.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
AnnData
|
AnnData object with RISE decomposition results. Must contain: - X.obsm["X_pf2_PaCMAP"]: PaCMAP embedding coordinates (n_cells, 2) - X.obsm["weighted_projections"]: Weighted cell projections (n_cells, rank) |
required |
cmp
|
int
|
Component number to visualize (1-indexed). For example, cmp=10 shows the cell associations for component 10. |
required |
ax
|
Axes
|
Matplotlib axes object to plot on. |
required |
cbarMax
|
float, optional (default: 1.0)
|
Maximum value for the color scale. Values are normalized to [-cbarMax, cbarMax]. Lower values increase contrast for components with weaker associations. |
1.0
|
Source code in scrise/plotting/pacmap.py
Rank Selection Plotting
scrise.plotting.rank_selection
Plotting functions for BiCV-based rank selection.
plot_bicv_r2x
Plot BiCV R2X and in-sample fit R2X across ranks.
The fit R2X (in-sample, computed on the full dataset) increases monotonically with rank. The BiCV R2X (held-out, averaged across repeated random cell/gene splits) penalizes overfitting and typically peaks near the rank that best generalizes to unseen data. The peak (or plateau) of the BiCV R2X curve is a good candidate for the rank to use.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
results
|
DataFrame
|
Output of :func: |
required |
ax
|
Axes
|
Matplotlib axes object to plot on. |
required |
Source code in scrise/plotting/rank_selection.py
Factor Stability
scrise.plotting.stability
Decomposition stability analysis and plotting functions.
plot_fms_diff_ranks
plot_fms_diff_ranks(
X: AnnData,
ax: Axes,
ranksList: list[int],
runs: int,
compress: int
| tuple[int, int | None]
| str
| bool
| None = "auto",
)
Plot Factor Match Score (FMS) across different ranks to assess stability.
FMS measures the reproducibility of PARAFAC2 decomposition results across multiple runs. Values above ~0.6 indicate stable, reproducible components. This helps determine which ranks produce reliable decompositions that are not overly sensitive to initialization or noise.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
AnnData
|
Preprocessed AnnData object containing single-cell RNA-seq data. Must have X.obs["condition_unique_idxs"] for RISE decomposition. |
required |
ax
|
Axes
|
Matplotlib axes object to plot on. |
required |
ranksList
|
list of int
|
List of rank values to test (e.g., [1, 5, 10, 15, 20, 25, 30]). Each rank will be run multiple times to compute FMS. |
required |
runs
|
int
|
Number of independent runs per rank to use for FMS calculation. Higher values give more reliable FMS estimates but take longer. Typical values: 3-5 runs. |
required |
compress
|
int | tuple[int, int | None] | str | bool | None
|
CANDELINC compression mode passed to |
'auto'
|
Notes
FMS values interpretation: - FMS > 0.9: Highly stable decomposition - FMS > 0.6: Acceptably stable decomposition - FMS < 0.6: Unstable, consider lower rank or more data
Source code in scrise/plotting/stability.py
Cell-Type Alignment Plotting
scrise.plotting.annotation_alignment
Plotting functions for cell-type alignment scoring.
plot_cell_type_alignment
plot_cell_type_alignment(
data: AnnData | CellTypeAlignmentResults | DataFrame,
ax: Axes | Sequence[Axes] | None = None,
cell_type_col: str = "cell_type",
projection_key: str = "weighted_projections",
signed: bool = False,
n_permutations: int = 1000,
alpha: float = 0.05,
annotate_significance: bool = True,
show_metrics: bool = True,
reorder: bool = True,
cmap=None,
metrics_cmap="Blues",
random_state=None,
) -> Axes | tuple[Axes, ...]
Plot cell-type alignment heatmap for RISE components.
Visualizes per-cell-type AUROC enrichment for each component with optional significance stars (* for q <= alpha) and row annotations for uniqueness (tau) and variance explained (eta^2).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
data
|
AnnData | CellTypeAlignmentResults | DataFrame
|
Alignment results or AnnData object containing RISE results. |
required |
ax
|
Axes | Sequence[Axes] | None
|
Matplotlib Axes to plot on. Can be: - None: creates a new figure with appropriate subplots. - Single Axes: plots the main AUROC heatmap on the provided axes. - Pair of Axes (ax_main, ax_metrics): plots heatmap on ax_main and metrics on ax_metrics. |
None
|
cell_type_col
|
str, optional (default: "cell_type")
|
Column name in data.obs containing cell type labels (if data is AnnData). |
'cell_type'
|
projection_key
|
str, optional (default: "weighted_projections")
|
Key in data.obsm containing projections (if data is AnnData). |
'weighted_projections'
|
signed
|
bool, optional (default: False)
|
If True, takes the absolute value of loadings. |
False
|
n_permutations
|
int, optional (default: 1000)
|
Number of permutations to run (if scoring an AnnData). |
1000
|
alpha
|
float, optional (default: 0.05)
|
FDR significance threshold. |
0.05
|
annotate_significance
|
bool, optional (default: True)
|
If True, marks significant cell types (q <= alpha, AUROC > 0.5) with an asterisk (*). |
True
|
show_metrics
|
bool, optional (default: True)
|
If True and subplots are available (or ax is None), shows tau and eta^2 row annotations. |
True
|
reorder
|
bool, optional (default: True)
|
If True, reorders components by dominant cell type. |
True
|
cmap
|
Colormap or str
|
Colormap for AUROC enrichment heatmap (defaults to blue-red diverging). |
None
|
metrics_cmap
|
Colormap or str, optional (default: "Blues")
|
Colormap for tau and eta^2 annotations. |
'Blues'
|
random_state
|
int | None
|
Random seed for permutations. |
None
|
Returns:
| Type | Description |
|---|---|
Axes | tuple[Axes, ...]
|
The plotted Matplotlib Axes. |
Source code in scrise/plotting/annotation_alignment.py
140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 | |