Ume fix perplexity device by karinazad · Pull Request #68 · prescient-design/lobster

karinazad · 2025-05-02T14:15:54Z

No description provided.

Copilot

Pull Request Overview

This PR addresses modifications in the UME model to change how perplexity metrics are stored and accessed during training and validation. Key changes include replacing the metrics dictionary with dynamic attributes using setattr, and adjusting the logging of perplexity metrics accordingly.

Comments suppressed due to low confidence (1)

src/lobster/model/_ume.py:146

[nitpick] Using attribute names with a slash (e.g., "train_perplexity/{modality.value}") is unconventional and may lead to confusion when accessing these attributes. Consider using a valid identifier format, such as replacing the slash with an underscore.

setattr(self, f"train_perplexity/{modality.value}", Perplexity(ignore_index=-100))

Copilot · 2025-05-02T18:41:14Z

src/lobster/model/_ume.py

+            metric = getattr(self, metric_name)
+            metric(logits_reshaped[mask], labels_reshaped[mask])
+
+            self.log(metric_name, metric, sync_dist=True)


The updated logging call no longer specifies the on_step parameter, which changes the logging behavior compared to previous implementations. If this change is intentional, please document the rationale to ensure consistent logging.

Suggested change

self.log(metric_name, metric, sync_dist=True)

self.log(metric_name, metric, on_step=True, sync_dist=True)

src/lobster/model/_ume.py

* pplx as attr * pplx as attr * pplx * comments * on step * comment

* peer fixes, add evaluate method * dataloader checkpoint callback (#60) * dataloader callback * utils * ume * gitignore dev * tests * update flash attention wheels (#61) * lock * torch 2.5 * torch 2.5 * part * .env * unpin flash attn (#62) * fix scheduler params (#64) * scheduler * fix scheduler * fix scheduler * Add AtomicaDataset (#63) Processed Atomica interactions dataset * Ume conversion/interaction tokenizer + fix SMILES and nucleotide tokenizers (#65) add two special tokens: <convert> and <interact> for later stages of Ume training: will be used as this: (or something like that) [CLS] PROT_SEQ [SEP] <convert> PROT_STRUCT(masked) [SEP] [CLS] PROT_SEQ [SEP] <interact> SMILES(masked) [SEP] extend functionality of UmeTokenizerTransform to handle dual modalities change the name of Ume embedding method and allow embedding from existing input_ids fix existing tokenizers: add lowercase normalized to nucleotide tokenizer (OG2 dataset contains a mix of upper and lowercase letters) BPE handled SMILES tokenization incorrectly, switch to WordLevel * Ume SMILES tokenizer fix (#66) * tokenizer * fix tests * lowercase normalizer for nt * tests * remove mod conv dataset * embed * Test * merge 2mod into UmeTokenizerTransform * fix tests * all * type hints * docstrings * tests * fix SMILES tokenizer * switch all tokenizer to BPE * Revert "switch all tokenizer to BPE" This reverts commit 367e77d. * tok * fix SMILES tokenizer * remove print statement * Ume perplexity logging (#67) * pplx * tests * src * ignore torchmetrics warnings * docstrings * docstrings * Update README.md (#69) * Ume fix perplexity device (#68) * pplx as attr * pplx as attr * pplx * comments * on step * comment * update tests, fix ruff * ruff * ruff ruff * Add <cls_modality> to Ume tokenizers (#71) * add <cls_modality> tokens * add <cls_modality> tokens * docstring * RNS metric implementation (#73) * add <cls_modality> tokens * add <cls_modality> tokens * modality embeddings * module dict * embeddings * tests * modality and device * rank zero only * rank zero * fix back modality mask * sync dist * RNS implementation * restore from main * restore * docstrings * docstrings * review * test * Ume modality-specific embeddings (#72) * add <cls_modality> tokens * add <cls_modality> tokens * modality embeddings * module dict * embeddings * tests * modality and device * rank zero only * rank zero * fix back modality mask * sync dist * add conversion transforms (#74) * add initial smiles to peptide and peptide to smiles transforms * remove smiles -> * transforms and touch up conversion functions * rename * add option to randomize smiles and caps --------- Co-authored-by: Colin Grambow <grambowc@gene.com> * fix def pad token, replace process_and_embed w/ ume.embed * update tests w -100 pad token --------- Co-authored-by: Taylor Joren <joren.taylor@gene.com> Co-authored-by: Karina Zadorozhny <karina.zadorozhny@gmail.com> Co-authored-by: Nathan Frey <ncfrey@users.noreply.github.com> Co-authored-by: Colin Grambow <17198155+cgrambow@users.noreply.github.com> Co-authored-by: Colin Grambow <grambowc@gene.com>

karinazad added 3 commits May 1, 2025 21:20

pplx as attr

be74c44

pplx as attr

b763931

pplx

d384c04

karinazad temporarily deployed to test.pypi.org May 2, 2025 14:29 — with GitHub Actions Inactive

ncfrey requested a review from Copilot May 2, 2025 18:40

Copilot AI reviewed May 2, 2025

View reviewed changes

ncfrey approved these changes May 2, 2025

View reviewed changes

ncfrey reviewed May 2, 2025

View reviewed changes

src/lobster/model/_ume.py Show resolved Hide resolved

comments

c7aec55

karinazad temporarily deployed to test.pypi.org May 2, 2025 18:48 — with GitHub Actions Inactive

on step

2b30322

karinazad temporarily deployed to test.pypi.org May 2, 2025 18:50 — with GitHub Actions Inactive

comment

3557077

karinazad temporarily deployed to test.pypi.org May 2, 2025 20:02 — with GitHub Actions Inactive

karinazad merged commit dc040c0 into main May 2, 2025
5 checks passed

karinazad deleted the ume-fix-perplexity-device branch May 2, 2025 20:17

taylormjs pushed a commit that referenced this pull request May 6, 2025

Ume fix perplexity device (#68)

a3ddc77

* pplx as attr * pplx as attr * pplx * comments * on step * comment

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Ume fix perplexity device#68

Ume fix perplexity device#68
karinazad merged 6 commits intomainfrom
ume-fix-perplexity-device

karinazad commented May 2, 2025

Uh oh!

Copilot AI left a comment

Uh oh!

Copilot AI May 2, 2025

Uh oh!

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

	self.log(metric_name, metric, sync_dist=True)
	self.log(metric_name, metric, on_step=True, sync_dist=True)

Conversation

karinazad commented May 2, 2025

Uh oh!

Copilot AI left a comment

Choose a reason for hiding this comment

Pull Request Overview

Uh oh!

Copilot AI May 2, 2025

Choose a reason for hiding this comment

Uh oh!

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants