{"urn":"urn:mavedb:00000661","publishedDate":"2023-11-13","id":716,"recordType":"ExperimentSet","experiments":[{"title":"COMT mid polysome abundance","shortDescription":"Integration of MAVEs to measure post-transcriptional gene expression for >3000 variants in COMT","abstractText":"We integrated three multiplexed assays to measure RNA abundance, protein abundance, and ribosome load for variants in the early coding region of the membrane-COMT (MB-COMT) protein isoform. Two regions that span clinically relevant variants (rs4633, rs4818, rs4680) in MB-COMT were targeted: codons 40-74 (region 1) and 136-158 (region 2). Our design enabled testing the importance of synonymous variation and a potential role of RNA secondary structure in controlling COMT gene expression.","methodText":"Variant library construction methods:\n\nPrimers used to generate the transgene are in Supplementary Table 4 located at https://github.com/ijhoskins/integrated_MAVEs/blob/main/supplementary/Supplementary_Tables.xlsx\n\nA plasmid containing the COMT noncoding exon 2 (NM_000754.4) and coding sequence was obtained from DNASU (clone HsCD00617865), and gBlocks were ordered for 1) the same COMT noncoding exon 2 with a downstream Flag tag; and 2) moxGFP (Costantini et al, 2015), which facilitates membrane localization. The 5′ UTR and coding sequence were amplified in Q5 polymerase reactions (NEB, M0491) with 10 ng HsCD00617865 template and 500 nM forward and reverse primers for 25 cycles using manufacturer’s recommendations with 55 ̊C annealing temperature. Products were purified with the Nucleospin PCR and Gel Cleanup Kit (Takara, 740609), and the 5′ UTR, CDS, and moxGFP gBlock were joined using a Gibson assembly reaction (NEB) following manufacturer recommendations.\n\n2 μL of the assembly reaction was used as template for a NEB Q5 PCR using Gateway-compatible primers with attB1/attB2 tails and the aforementioned PCR conditions. The product was subsequently transferred into a pDONR223 entry clone via a Gateway BP reaction (Thermo Fisher Scientific), and transferred to a destination vector (pDEST_HC_Rec_Bxb_v2) compatible with recombination into the 293T LLP iCasp9 Blast cell line with a 1 h LR reaction (Thermo Fisher Scientific, 11-791-020). Plasmid maps for pDONR223 and pDEST_HC_Rec_Bxb_v2 are located at https://github.com/ijhoskins/integrated_MAVEs/tree/main/plasmid_maps. The isolated clone had a deletion in the 5′ UTR, so a new clone was generated using the deletion clone as a template.\n\nA CDS-moxGFP product was amplified from the deletion clone in a 50 μL NEB Q5 reaction for 25 cycles with 30 s annealing at 55 ̊C. The product was extracted from a 1% agarose-TBE gel (Zymo Gel Extraction kit, 10 μL elution), then assembled with the 5′ UTR gBlock in a Gibson reaction. A final 50 μL Q5 PCR reaction was performed with 2 μL assembly reaction template for 30 cycles with 30 s annealing at 50 ̊C. The product was extracted from a 1% agarose-TBE gel as previously detailed. The transgene was transferred to the pDONR223 and pDEST_HC_Rec_Bxb_v2 with 1 h Gateway BP and LR reactions, respectively. The isolated transgene contained the full 91 nt of exon 2 of the MB-COMT 5′ UTR, a N-terminal Flag tag, a C-terminal moxGFP tag upstream of a bicistronic IRES-mCherry element, and partial vector-derived 5′ and 3′ UTRs from the landing pad vector pDEST_HC_Rec_Bxb_v2.\n\nThe 5′ UTR-Flag-COMT-moxGFP clone in pDEST_HC_Rec_Bxb_v2 was sent to Twist Biosciences for mutagenesis of membrane-COMT codons 40-74 (region 1) and 136-158 (region 2) using degenerate primers. All silent and missense variants and the amber nonsense variant were generated at each codon in the template with added attB recombination sites. Inserts for each target region (region 1, region 2) were received and pooled separately, with codons in each region pooled at equimolar ratios. A Gateway BP reaction was performed with 70 ng of pooled inserts and 150 ng pDONR223 at 25 ̊C overnight (19 h). 1.5 μL BP reaction was electroporated into 25 μL Endura Electrocompetent cells (Lucigen, 60242) with the following conditions: 2 mm cuvette, 2500 V, 200 Ohms, 25 μF. Cells were outgrown for 1 h at 37 ̊C in 1 mL Lucigen outgrowth media, then 600 μL was plated on Nunc Square Bioassay dishes, scraped in 8 mL LB Miller broth, and plasmid library purified with the ZymoPURE II Plasmid Maxiprep Kit (Zymo Research, D4203) using 200 μL 50 ̊C elution buffer, and including endotoxin removal. Library sizes were estimated at 942,000 and 510,000 clones for region 1 and region 2 libraries, respectively (or 434-fold and 357-fold coverage of each variant). 1 μg of each library in pDONR223 was digested with Blp1 (NEB) for 1 h 15 min at 37 ̊C. Digests were run on a 0.8% agarose-TAE gel and full-length plasmids were extracted with the Nucleospin PCR and Gel Cleanup Kit (Takara, 740609). A LR reaction was performed with 150 ng entry library and 150 ng pDEST_HC_Rec_Bxb_v2, incubated at 25 ̊C for 21 h. 25 μL Endura Electrocompetent cells were transformed with 1.5 μL LR reaction using the same conditions as for the BP reaction. After 14.5 h, colonies were collected with 7 mL LB Miller broth, and 3 mL was pelleted and processed with the ZymoPURE II Plasmid Maxiprep Kit.\n\nDescription of functional assay, including model system and selection type:\n\nFor each region 1 and region 2 biological replicate cell line, 20 μg of the library along with an equal mass of Bxb1 recombinase (pCAG-NLS-HA-Bxb1) was transfected into a 15 cm dish of HEK293T LLP iCasp9 Blast cells using Lipofectamine 3000 (Thermo Fisher Scientific, L3000008), with volumes scaled based on 3.75 μL reagent per 6-well. After 48 h, at near full confluency, 2 μg/mL Doxycycline and 10 nM AP1903 (MedChemExpress, HY-16046), both solubilized in DMSO, were added for negative selection of non-recombined cells. The next day, dead cells were removed and recombined cells were grown out for an additional two days with fresh media containing Doxycycline and AP1903. Cells were recovered for two days by growth in media without Doxycycline and AP1903. Before functional assays, transcription was induced with 2 μg/mL Doxycycline for 24 h (total RNA, polysome RNA readouts) or 21 h (flow cytometry readouts), and cells were stimulated with fresh media for 1 h 15 min prior to harvest. Importantly, because the COMT transgene is membrane localized and faces the extracellular space, cells are collected by pipetting in cold PBS, and not by trypsinization.\n\nPolysome fractionation and polysomal RNA purification\nFor COMT library stable cell lines, cells at 80-95% confluency (one 15 cm dish per gradient) were incubated with 100 μg/mL cycloheximide (Sigma Aldrich, C4859) in media for 10 min at 37 ̊C, then resuspended in ice cold PBS with 100 μg/mL cycloheximide (PBS-CHX), collected at 4 ̊C by pipetting, and pelleted at 300 x g for 7 min at 4 ̊C. Pellets were flash frozen in liquid nitrogen, then lysed on ice for 10 min with 300-400 μL lysis buffer (20 mM Tris-HCl pH 7.5; 150 mM NaCl; 5 mM MgCl2, 1 mM fresh DTT, 20 mM ribonucleoside vanadyl complex (Fisher Scientific, 50-812-650), 1X protease inhibitor cocktail EDTA-free (Fisher Scientific, 539196), and 100 μg/ml cycloheximide. Lysates were clarified by centrifugation at 1,300 x g for 10 min at 4 ̊C, and 1/10th of the lysate was saved for total RNA extraction. 20-50% sucrose gradients were made with the BioComp Gradient Maker in 20 mM Tris-HCl pH 7.5, 150 mM NaCl, 5 mM MgCl2, 1 mM DTT, and 100 μg/ml cycloheximide using Beckman Coulter UC tubes 9/16 x 3-1/2 (Fisher Scientific, NC9194790), and equilibrating overnight at 4 ̊C. Clarified lysates were loaded onto sucrose gradients, and fractionated by ultracentrifugation at 38,000 rpm for 2.5 h at 4 ̊C in a Beckman Coulter SW-41 Ti rotor. Approximately 500 μL polysome fractions were collected using the Biocomp Piston Gradient Fractionator (v8.04) with a Triax flow cell (BioComp model FC-1) and Gilson Fraction Collector using a piston speed of 0.2 mm/s. RNA was precipitated from raw polysome fractions by addition of sodium acetate to 300 mM pH 5.2, 2 μL Glycoblue, and 2x volumes cold ethanol, stored overnight at -20  ̊C. RNA was resuspended in 50-80 μL ultrapure water, and these raw fractions were pooled into metafractions and extracted by phenol-chloroform (5:1) and ethanol precipitation overnight at -20 ̊C. After a single 70% ethanol wash, polysomal RNA was resuspended in 30 μL ultrapure water. Lastly, DNaseI treatment and a column purification was performed prior to library preparation.\n\nFlow cytometry analysis and flow cytometry sorting\nFollowing induction with Doxycycline and media stimulation, cells were washed once with cold PBS, resuspended with PBS, and assayed in a 5 mL Falcon polystyrene round bottom tube with a cell-strainer cap (Fisher Scientific, 08-771-23). Analysis of point mutants and deletion cell lines was conducted on a LSRFortessa SORP instrument (BD) with FACSDiva v6.1.3, and flow cytometry was conducted on a FACSAria Fusion SORP instrument (BD) at the Center for Biomedical Research Support Microscopy and Imaging Facility at UT Austin (RRID# SCR_021756). Gating was performed with FlowJo. A gate for cells versus debris was defined based on FSC-H versus SSC-H, and another gate for single cells versus doublets was defined based on FSC-A versus FSC-H. Fluorescence gating was based on autofluorescence of 239T LLP iCasp9 Blast cells without recombination or of a stable cell line expressing the template transgene but without Doxycycline induction. After flow cytometry sorting, each population was expanded up to a confluent T-25 flask, washed, trypsinized, pelleted, and flash frozen prior to gDNA extraction.\n\n\nSequencing strategy and sequencing technology:\n\ngDNA extraction\ngDNA was extracted from approximately 3-4 million cells with the Cell and Tissue DNA Isolation Kit (Norgen Biotek Corp, 24700), including RNaseA treatment at 37 ̊C for 15 min, and eluting in 200 μL warm elution buffer.\n\nRNA extraction and cDNA synthesis\nApproximately 3-4 million cells were solubilized with 1 mL QIAzol (Qiagen, 79306) and 0.2 mL chloroform in 5PRIME Phase-Lock Gel heavy tubes (QuantaBio, 2302830), according to the manufacturer’s recommendations. RNA was precipitated at -20 ̊C following the addition of 2 μL GlycoBlue (Thermo Fisher, AM9515) and 2.5 volumes of cold absolute ethanol. RNA was washed once with cold 70% ethanol then resuspended in 30 μL water. 10 μg total RNA or 15 μL of polysomal RNA (half of 30 μL eluate, volumetric inputs) was treated with DNaseI (NEB, M0303) in a 100 μL reaction at 37 ̊C for 15 min, then re-purified by the RNA Clean and Concentrator Kit (Zymo Research, R1015) and eluted in 20 μL water. 10 μL DNaseI-treated total RNA (<10 μg) or DNaseI-treated polysomal RNA (200 ng - 5 μg) was denatured at 65 ̊C for 5 min followed by RT primer annealing at 4 ̊C for 2 min, using 2 pmol pDEST_HC_Rec_Bxb_v2_R primer specific for the landing pad (see Supplementary Table 4) and 2.5 mM random hexamers. Primed total RNA was included in 20 μL SuperScript IV cDNA synthesis reaction (Thermo Fisher, 18090010) with SUPERase-In RNase inhibitor (Thermo Fisher, AM2696), and first-strand cDNA was synthesized by incubating at 23 ̊C for 10 min, 55 ̊C for 1 h, followed by RT inactivation at 80 ̊C for 10 min. RNA was digested with addition of 5 U RNaseH (NEB, M0297) to the first strand cDNA synthesis reaction and incubation at 37 ̊C for 20 min.\n \nLanding pad amplification (PCR1) \nNucleic acids (gDNA, 1st strand cDNA) were amplified with Q5 polymerase (NEB, M0491) in 50 μL PCR reactions for 16 cycles with 500 nM landing-pad-specific primers (pDEST_HC_Rec_Bxb_v2_F/R) flanking the entire COMT insert (~1.7 kb). For negative control (plasmid) samples, 10 ng was used as template. For total RNA and polysome RNA samples, 15 μL of unpurified 1st strand cDNA synthesis reaction was used as template. For total gDNA samples, 4 μg template was used in each of three PCR reactions. For flow-sorted gDNA samples, either 1.75 μg template was used in three PCR reactions or 2.5 μg was input into two reactions, with 14 cycles instead of 16. The cycling parameters were: initial denaturation at 98 ̊C for 30 s; 3-step cycling with denaturation at 98 ̊C for 10 s, anneal at 65 ̊C for 30 s, extension at 72 ̊C for 1 min; final extension at 72 ̊C for 2 min. Products from total RNA and polysome RNA samples were purified with the QIAquick PCR Purification Kit (Qiagen, 28106) and eluted in 30 μL elution buffer. Products for total gDNA samples were pooled, resolved on a 0.8% agarose/TAE gel, stained with 1x SYBR Gold, and extracted using the Nucleospin Gel and PCR Cleanup Kit (Takara, 740609) or Zymoclean Gel DNA Recovery Kit (Zymo Research, D4007) with 20 μL 70 ̊C elution buffer. At this point, PCR1 products for gDNA samples cannot typically be seen in the gel, and the expected band size is cut using comparison to a high-molecular weight ladder.\n\nCoding sequence amplification (PCR2) \nPurified PCR1 products (landing pad insert) were amplified for each of COMT target regions in a 50 μL NEB Q5 reaction (NEB, M0491) for 8 cycles, following the same cycling parameters as for PCR1 and 500 nM tailed primers (see Supplementary Table 4). 15 μL PCR1 product was used as template (50% of eluate for total RNA and polysome RNA samples; 75% of eluate for gel-extracted gDNA samples). Products were purified with the Nucleospin Gel and PCR Cleanup Kit (Takara, 740609) with 1:5 buffer NTI dilution and 25 μL 70 ̊C elution buffer, or with the QIAquick PCR Purification Kit (Qiagen, 28106) with 30 μL elution buffer.\n\nIllumina adapter addition (PCR3) \nA final Q5 PCR reaction was carried out for 8 cycles with the same formulation as PCR2 but using NEBNext Multiplex Oligos for Illumina Dual Index Primers Set 1 (NEB, E7600S) according to manufacturer’s recommendations (65 ̊C annealing). 25 μL/30 μL eluate was used as template for total gDNA, total RNA, and polysome RNA samples. For flow cytometry gDNA samples, 15 μL/25 μL eluate was used as template. PCR3 products (final library) were purified with the Nucleospin Gel and PCR Cleanup kit (Takara, 740609) and eluted in 25 μL 70 ̊C buffer.\n\nQuantification and size-selection of final libraries \nFinal libraries were quantified using the KAPA Library Quantification Kit for Illumina (Roche, KK4873) according to the manufacturer’s recommendations, using 1:10,000 dilution of libraries and size-correction with an average fragment size of 300 nt (region 1) and 285 nt (region 2). qPCR was performed on the Applied Biosystems ViiA 7 instrument and values were used to pool individual libraries at equimolar ratios to target ~5 M read pairs per library. Final library pools were size-selected use PAGE purification on a 15% TBE gel followed by a crush-and-soak method (Sambrook & Russell, 2006), or with BluePippin 2% gel (Sage Science, BDF2010) to select fragments from 250-340 bp. Following size-selection, final pools were quantified by High Sensitivity DNA Kit (Agilent, 5067-4626) prior to sequencing.\n\nNext-Generation Sequencing\nAll libraries were sequenced 2 x 150 bp (paired-end) on Illumina platforms with minimum 5% PhiX. Libraries were sequenced at MedGenome, Inc. on the HiSeq X or at Novogene Corporation, Inc. on the NovaSeq 6000. Raw FASTQs were obtained directly from the vendor.\n\nStructure of biological and technical replicates:\nFour biological replicate cell lines for each target ROI were generated, with some biological replicates having technical replication (replication of growth, library preparation, and sequencing). The median variant frequency of technical replicates was taken prior to statistical testing with ALDEx2 as described below. \n\n\nSequencing read filtering approach:\n\nVariants were called with satmut_utils v1.0.3-dev001 (Hoskins, 2023). Analysis of COMT libraries prepared with the amplicon method utilized the satmut_utils call parameters `-n 2 -m 1 --r1_fiveprime_adapters TACACGACGCTCTTCCGATCT --r1_threeprime_adapters AGATCGGAAGAGCACACGTCT --r2_fiveprime_adapters AGACGTGTGCTCTTCCGATCT --r2_threeprime_adapters AGATCGGAAGAGCGTCGTGTA`.\n\nRandom forest models were trained on simulated data as previously described for libraries prepared from the amplicon method (Hoskins et al, 2023). Variants were post-processed by the following criteria: variants are within the mutagenesis regions; 2) nonsense variants must match the amber codon UAG; 3) log10 variant frequency is > -5.8; 4) SNP variants with false positive random forest predictions (p > 0.5) in more than half of the gDNA libraries were filtered out. Amino acid positions were adjusted to account for the Flag tag such that positions align with the endogenous membrane-COMT protein isoform. \n\n\nDescription of statistical model for converting counts to scores, including normalization:\n\nFrequencies were normalized to the wild-type count at each position (variant count + 0.5 / wild-type count + 0.5) (Rubin et al, 2017). For each variant, normalized frequencies in the negative control (plasmid template) were subtracted from the frequency observed in total gDNA, total RNA, polysomal RNA, and flow cytometry gDNA samples, and variants with resultant negative values were filtered out. Normalized frequencies were batch-corrected to account for the library preparation experiment using ComBat (Leek et al, 2022) with default parameters. Total gDNA and total RNA frequencies were adjusted with a model matrix, whereas no model matrix was used for polysomal RNA and flow cytometry gDNA libraries, as readouts to be compared were in different batches. The median among technical replicates was computed for each biological replicate (independent stable cell line).\n\nWe only considered variants that were observed in all biological replicates for each readout to be compared such that median log-frequency after wild-type normalization and batch correction was greater than -10. These frequencies were converted to counts per million and input to ALDEx2 (Fernandes et al, 2013, 2014; Gloor et al, 2016; Gloor, 2021) functions aldex.clr, aldex.effect, and aldex.ttest with parameters `paired=FALSE, denom=”iqlr”, mc.samples=128`.\n\nALDEx2 effect size is an estimate of the median standardized difference between groups. To avoid any distributional assumptions for standardization, ALDEx2 uses a permutation based non-parametric estimate of dispersion. We opted to use ALDEx2 75% confidence intervals instead of reporting standard deviations on fold changes. The confidence intervals are determined by Monte Carlo methods that produce a posterior probability distribution of the observed data given repeated sampling. The comparisons made were 1) total RNA to total gDNA; 2) polysome metafraction 3 or 4 to total RNA; 3) flow cytometry population 3 (low moxGFP fluorescence compared to mCherry fluorescence) to total gDNA.\n\n\nDescription of additional data columns included in the score or count tables, including column naming conventions:\n\nscore: ALDEx2 effect size estimate.\nscore_ci_lower: lower 75% confidence interval of ALDEx2 effect size estimate.\nscore_ci_higher: higher 75% confidence interval of ALDEx2 effect size estimate.\nscore_FDR: false discovery rate of ALDEx2 effect size estimate.\n","extraMetadata":{},"recordType":"Experiment","urn":"urn:mavedb:00000661-a","createdBy":{"orcidId":"0000-0003-2307-3299","firstName":"Ian","lastName":"Hoskins","recordType":"User"},"modifiedBy":{"orcidId":"0000-0003-2307-3299","firstName":"Ian","lastName":"Hoskins","recordType":"User"},"creationDate":"2023-11-13","modificationDate":"2023-11-13","publishedDate":"2023-11-13","experimentSetUrn":"urn:mavedb:00000661","doiIdentifiers":[{"identifier":"10.1101/2023.08.02.551517","id":57,"recordType":"DoiIdentifier","url":"https://doi.org/10.1101/2023.08.02.551517"}],"primaryPublicationIdentifiers":[],"secondaryPublicationIdentifiers":[],"rawReadIdentifiers":[{"identifier":"GSE246139","id":87,"recordType":"RawReadIdentifier","url":"http://www.ebi.ac.uk/ena/data/view/GSE246139"}],"contributors":[],"keywords":[],"scoreSetUrns":["urn:mavedb:00000661-a-1"],"externalLinks":{},"numScoreSets":1,"processingState":null,"officialCollections":[]},{"title":"COMT heavy polysome abundance","shortDescription":"Integration of MAVEs to measure post-transcriptional gene expression for >3000 variants in COMT","abstractText":"We integrated three multiplexed assays to measure RNA abundance, protein abundance, and ribosome load for variants in the early coding region of the membrane-COMT (MB-COMT) protein isoform. Two regions that span clinically relevant variants (rs4633, rs4818, rs4680) in MB-COMT were targeted: codons 40-74 (region 1) and 136-158 (region 2). Our design enabled testing the importance of synonymous variation and a potential role of RNA secondary structure in controlling COMT gene expression.","methodText":"Variant library construction methods:\n\nPrimers used to generate the transgene are in Supplementary Table 4 located at https://github.com/ijhoskins/integrated_MAVEs/blob/main/supplementary/Supplementary_Tables.xlsx\n\nA plasmid containing the COMT noncoding exon 2 (NM_000754.4) and coding sequence was obtained from DNASU (clone HsCD00617865), and gBlocks were ordered for 1) the same COMT noncoding exon 2 with a downstream Flag tag; and 2) moxGFP (Costantini et al, 2015), which facilitates membrane localization. The 5′ UTR and coding sequence were amplified in Q5 polymerase reactions (NEB, M0491) with 10 ng HsCD00617865 template and 500 nM forward and reverse primers for 25 cycles using manufacturer’s recommendations with 55 ̊C annealing temperature. Products were purified with the Nucleospin PCR and Gel Cleanup Kit (Takara, 740609), and the 5′ UTR, CDS, and moxGFP gBlock were joined using a Gibson assembly reaction (NEB) following manufacturer recommendations.\n\n2 μL of the assembly reaction was used as template for a NEB Q5 PCR using Gateway-compatible primers with attB1/attB2 tails and the aforementioned PCR conditions. The product was subsequently transferred into a pDONR223 entry clone via a Gateway BP reaction (Thermo Fisher Scientific), and transferred to a destination vector (pDEST_HC_Rec_Bxb_v2) compatible with recombination into the 293T LLP iCasp9 Blast cell line with a 1 h LR reaction (Thermo Fisher Scientific, 11-791-020). Plasmid maps for pDONR223 and pDEST_HC_Rec_Bxb_v2 are located at https://github.com/ijhoskins/integrated_MAVEs/tree/main/plasmid_maps. The isolated clone had a deletion in the 5′ UTR, so a new clone was generated using the deletion clone as a template.\n\nA CDS-moxGFP product was amplified from the deletion clone in a 50 μL NEB Q5 reaction for 25 cycles with 30 s annealing at 55 ̊C. The product was extracted from a 1% agarose-TBE gel (Zymo Gel Extraction kit, 10 μL elution), then assembled with the 5′ UTR gBlock in a Gibson reaction. A final 50 μL Q5 PCR reaction was performed with 2 μL assembly reaction template for 30 cycles with 30 s annealing at 50 ̊C. The product was extracted from a 1% agarose-TBE gel as previously detailed. The transgene was transferred to the pDONR223 and pDEST_HC_Rec_Bxb_v2 with 1 h Gateway BP and LR reactions, respectively. The isolated transgene contained the full 91 nt of exon 2 of the MB-COMT 5′ UTR, a N-terminal Flag tag, a C-terminal moxGFP tag upstream of a bicistronic IRES-mCherry element, and partial vector-derived 5′ and 3′ UTRs from the landing pad vector pDEST_HC_Rec_Bxb_v2.\n\nThe 5′ UTR-Flag-COMT-moxGFP clone in pDEST_HC_Rec_Bxb_v2 was sent to Twist Biosciences for mutagenesis of membrane-COMT codons 40-74 (region 1) and 136-158 (region 2) using degenerate primers. All silent and missense variants and the amber nonsense variant were generated at each codon in the template with added attB recombination sites. Inserts for each target region (region 1, region 2) were received and pooled separately, with codons in each region pooled at equimolar ratios. A Gateway BP reaction was performed with 70 ng of pooled inserts and 150 ng pDONR223 at 25 ̊C overnight (19 h). 1.5 μL BP reaction was electroporated into 25 μL Endura Electrocompetent cells (Lucigen, 60242) with the following conditions: 2 mm cuvette, 2500 V, 200 Ohms, 25 μF. Cells were outgrown for 1 h at 37 ̊C in 1 mL Lucigen outgrowth media, then 600 μL was plated on Nunc Square Bioassay dishes, scraped in 8 mL LB Miller broth, and plasmid library purified with the ZymoPURE II Plasmid Maxiprep Kit (Zymo Research, D4203) using 200 μL 50 ̊C elution buffer, and including endotoxin removal. Library sizes were estimated at 942,000 and 510,000 clones for region 1 and region 2 libraries, respectively (or 434-fold and 357-fold coverage of each variant). 1 μg of each library in pDONR223 was digested with Blp1 (NEB) for 1 h 15 min at 37 ̊C. Digests were run on a 0.8% agarose-TAE gel and full-length plasmids were extracted with the Nucleospin PCR and Gel Cleanup Kit (Takara, 740609). A LR reaction was performed with 150 ng entry library and 150 ng pDEST_HC_Rec_Bxb_v2, incubated at 25 ̊C for 21 h. 25 μL Endura Electrocompetent cells were transformed with 1.5 μL LR reaction using the same conditions as for the BP reaction. After 14.5 h, colonies were collected with 7 mL LB Miller broth, and 3 mL was pelleted and processed with the ZymoPURE II Plasmid Maxiprep Kit.\n\nDescription of functional assay, including model system and selection type:\n\nFor each region 1 and region 2 biological replicate cell line, 20 μg of the library along with an equal mass of Bxb1 recombinase (pCAG-NLS-HA-Bxb1) was transfected into a 15 cm dish of HEK293T LLP iCasp9 Blast cells using Lipofectamine 3000 (Thermo Fisher Scientific, L3000008), with volumes scaled based on 3.75 μL reagent per 6-well. After 48 h, at near full confluency, 2 μg/mL Doxycycline and 10 nM AP1903 (MedChemExpress, HY-16046), both solubilized in DMSO, were added for negative selection of non-recombined cells. The next day, dead cells were removed and recombined cells were grown out for an additional two days with fresh media containing Doxycycline and AP1903. Cells were recovered for two days by growth in media without Doxycycline and AP1903. Before functional assays, transcription was induced with 2 μg/mL Doxycycline for 24 h (total RNA, polysome RNA readouts) or 21 h (flow cytometry readouts), and cells were stimulated with fresh media for 1 h 15 min prior to harvest. Importantly, because the COMT transgene is membrane localized and faces the extracellular space, cells are collected by pipetting in cold PBS, and not by trypsinization.\n\nPolysome fractionation and polysomal RNA purification\nFor COMT library stable cell lines, cells at 80-95% confluency (one 15 cm dish per gradient) were incubated with 100 μg/mL cycloheximide (Sigma Aldrich, C4859) in media for 10 min at 37 ̊C, then resuspended in ice cold PBS with 100 μg/mL cycloheximide (PBS-CHX), collected at 4 ̊C by pipetting, and pelleted at 300 x g for 7 min at 4 ̊C. Pellets were flash frozen in liquid nitrogen, then lysed on ice for 10 min with 300-400 μL lysis buffer (20 mM Tris-HCl pH 7.5; 150 mM NaCl; 5 mM MgCl2, 1 mM fresh DTT, 20 mM ribonucleoside vanadyl complex (Fisher Scientific, 50-812-650), 1X protease inhibitor cocktail EDTA-free (Fisher Scientific, 539196), and 100 μg/ml cycloheximide. Lysates were clarified by centrifugation at 1,300 x g for 10 min at 4 ̊C, and 1/10th of the lysate was saved for total RNA extraction. 20-50% sucrose gradients were made with the BioComp Gradient Maker in 20 mM Tris-HCl pH 7.5, 150 mM NaCl, 5 mM MgCl2, 1 mM DTT, and 100 μg/ml cycloheximide using Beckman Coulter UC tubes 9/16 x 3-1/2 (Fisher Scientific, NC9194790), and equilibrating overnight at 4 ̊C. Clarified lysates were loaded onto sucrose gradients, and fractionated by ultracentrifugation at 38,000 rpm for 2.5 h at 4 ̊C in a Beckman Coulter SW-41 Ti rotor. Approximately 500 μL polysome fractions were collected using the Biocomp Piston Gradient Fractionator (v8.04) with a Triax flow cell (BioComp model FC-1) and Gilson Fraction Collector using a piston speed of 0.2 mm/s. RNA was precipitated from raw polysome fractions by addition of sodium acetate to 300 mM pH 5.2, 2 μL Glycoblue, and 2x volumes cold ethanol, stored overnight at -20  ̊C. RNA was resuspended in 50-80 μL ultrapure water, and these raw fractions were pooled into metafractions and extracted by phenol-chloroform (5:1) and ethanol precipitation overnight at -20 ̊C. After a single 70% ethanol wash, polysomal RNA was resuspended in 30 μL ultrapure water. Lastly, DNaseI treatment and a column purification was performed prior to library preparation.\n\nFlow cytometry analysis and flow cytometry sorting\nFollowing induction with Doxycycline and media stimulation, cells were washed once with cold PBS, resuspended with PBS, and assayed in a 5 mL Falcon polystyrene round bottom tube with a cell-strainer cap (Fisher Scientific, 08-771-23). Analysis of point mutants and deletion cell lines was conducted on a LSRFortessa SORP instrument (BD) with FACSDiva v6.1.3, and flow cytometry was conducted on a FACSAria Fusion SORP instrument (BD) at the Center for Biomedical Research Support Microscopy and Imaging Facility at UT Austin (RRID# SCR_021756). Gating was performed with FlowJo. A gate for cells versus debris was defined based on FSC-H versus SSC-H, and another gate for single cells versus doublets was defined based on FSC-A versus FSC-H. Fluorescence gating was based on autofluorescence of 239T LLP iCasp9 Blast cells without recombination or of a stable cell line expressing the template transgene but without Doxycycline induction. After flow cytometry sorting, each population was expanded up to a confluent T-25 flask, washed, trypsinized, pelleted, and flash frozen prior to gDNA extraction.\n\n\nSequencing strategy and sequencing technology:\n\ngDNA extraction\ngDNA was extracted from approximately 3-4 million cells with the Cell and Tissue DNA Isolation Kit (Norgen Biotek Corp, 24700), including RNaseA treatment at 37 ̊C for 15 min, and eluting in 200 μL warm elution buffer.\n\nRNA extraction and cDNA synthesis\nApproximately 3-4 million cells were solubilized with 1 mL QIAzol (Qiagen, 79306) and 0.2 mL chloroform in 5PRIME Phase-Lock Gel heavy tubes (QuantaBio, 2302830), according to the manufacturer’s recommendations. RNA was precipitated at -20 ̊C following the addition of 2 μL GlycoBlue (Thermo Fisher, AM9515) and 2.5 volumes of cold absolute ethanol. RNA was washed once with cold 70% ethanol then resuspended in 30 μL water. 10 μg total RNA or 15 μL of polysomal RNA (half of 30 μL eluate, volumetric inputs) was treated with DNaseI (NEB, M0303) in a 100 μL reaction at 37 ̊C for 15 min, then re-purified by the RNA Clean and Concentrator Kit (Zymo Research, R1015) and eluted in 20 μL water. 10 μL DNaseI-treated total RNA (<10 μg) or DNaseI-treated polysomal RNA (200 ng - 5 μg) was denatured at 65 ̊C for 5 min followed by RT primer annealing at 4 ̊C for 2 min, using 2 pmol pDEST_HC_Rec_Bxb_v2_R primer specific for the landing pad (see Supplementary Table 4) and 2.5 mM random hexamers. Primed total RNA was included in 20 μL SuperScript IV cDNA synthesis reaction (Thermo Fisher, 18090010) with SUPERase-In RNase inhibitor (Thermo Fisher, AM2696), and first-strand cDNA was synthesized by incubating at 23 ̊C for 10 min, 55 ̊C for 1 h, followed by RT inactivation at 80 ̊C for 10 min. RNA was digested with addition of 5 U RNaseH (NEB, M0297) to the first strand cDNA synthesis reaction and incubation at 37 ̊C for 20 min.\n \nLanding pad amplification (PCR1) \nNucleic acids (gDNA, 1st strand cDNA) were amplified with Q5 polymerase (NEB, M0491) in 50 μL PCR reactions for 16 cycles with 500 nM landing-pad-specific primers (pDEST_HC_Rec_Bxb_v2_F/R) flanking the entire COMT insert (~1.7 kb). For negative control (plasmid) samples, 10 ng was used as template. For total RNA and polysome RNA samples, 15 μL of unpurified 1st strand cDNA synthesis reaction was used as template. For total gDNA samples, 4 μg template was used in each of three PCR reactions. For flow-sorted gDNA samples, either 1.75 μg template was used in three PCR reactions or 2.5 μg was input into two reactions, with 14 cycles instead of 16. The cycling parameters were: initial denaturation at 98 ̊C for 30 s; 3-step cycling with denaturation at 98 ̊C for 10 s, anneal at 65 ̊C for 30 s, extension at 72 ̊C for 1 min; final extension at 72 ̊C for 2 min. Products from total RNA and polysome RNA samples were purified with the QIAquick PCR Purification Kit (Qiagen, 28106) and eluted in 30 μL elution buffer. Products for total gDNA samples were pooled, resolved on a 0.8% agarose/TAE gel, stained with 1x SYBR Gold, and extracted using the Nucleospin Gel and PCR Cleanup Kit (Takara, 740609) or Zymoclean Gel DNA Recovery Kit (Zymo Research, D4007) with 20 μL 70 ̊C elution buffer. At this point, PCR1 products for gDNA samples cannot typically be seen in the gel, and the expected band size is cut using comparison to a high-molecular weight ladder.\n\nCoding sequence amplification (PCR2) \nPurified PCR1 products (landing pad insert) were amplified for each of COMT target regions in a 50 μL NEB Q5 reaction (NEB, M0491) for 8 cycles, following the same cycling parameters as for PCR1 and 500 nM tailed primers (see Supplementary Table 4). 15 μL PCR1 product was used as template (50% of eluate for total RNA and polysome RNA samples; 75% of eluate for gel-extracted gDNA samples). Products were purified with the Nucleospin Gel and PCR Cleanup Kit (Takara, 740609) with 1:5 buffer NTI dilution and 25 μL 70 ̊C elution buffer, or with the QIAquick PCR Purification Kit (Qiagen, 28106) with 30 μL elution buffer.\n\nIllumina adapter addition (PCR3) \nA final Q5 PCR reaction was carried out for 8 cycles with the same formulation as PCR2 but using NEBNext Multiplex Oligos for Illumina Dual Index Primers Set 1 (NEB, E7600S) according to manufacturer’s recommendations (65 ̊C annealing). 25 μL/30 μL eluate was used as template for total gDNA, total RNA, and polysome RNA samples. For flow cytometry gDNA samples, 15 μL/25 μL eluate was used as template. PCR3 products (final library) were purified with the Nucleospin Gel and PCR Cleanup kit (Takara, 740609) and eluted in 25 μL 70 ̊C buffer.\n\nQuantification and size-selection of final libraries \nFinal libraries were quantified using the KAPA Library Quantification Kit for Illumina (Roche, KK4873) according to the manufacturer’s recommendations, using 1:10,000 dilution of libraries and size-correction with an average fragment size of 300 nt (region 1) and 285 nt (region 2). qPCR was performed on the Applied Biosystems ViiA 7 instrument and values were used to pool individual libraries at equimolar ratios to target ~5 M read pairs per library. Final library pools were size-selected use PAGE purification on a 15% TBE gel followed by a crush-and-soak method (Sambrook & Russell, 2006), or with BluePippin 2% gel (Sage Science, BDF2010) to select fragments from 250-340 bp. Following size-selection, final pools were quantified by High Sensitivity DNA Kit (Agilent, 5067-4626) prior to sequencing.\n\nNext-Generation Sequencing\nAll libraries were sequenced 2 x 150 bp (paired-end) on Illumina platforms with minimum 5% PhiX. Libraries were sequenced at MedGenome, Inc. on the HiSeq X or at Novogene Corporation, Inc. on the NovaSeq 6000. Raw FASTQs were obtained directly from the vendor.\n\nStructure of biological and technical replicates:\nFour biological replicate cell lines for each target ROI were generated, with some biological replicates having technical replication (replication of growth, library preparation, and sequencing). The median variant frequency of technical replicates was taken prior to statistical testing with ALDEx2 as described below. \n\n\nSequencing read filtering approach:\n\nVariants were called with satmut_utils v1.0.3-dev001 (Hoskins, 2023). Analysis of COMT libraries prepared with the amplicon method utilized the satmut_utils call parameters `-n 2 -m 1 --r1_fiveprime_adapters TACACGACGCTCTTCCGATCT --r1_threeprime_adapters AGATCGGAAGAGCACACGTCT --r2_fiveprime_adapters AGACGTGTGCTCTTCCGATCT --r2_threeprime_adapters AGATCGGAAGAGCGTCGTGTA`.\n\nRandom forest models were trained on simulated data as previously described for libraries prepared from the amplicon method (Hoskins et al, 2023). Variants were post-processed by the following criteria: variants are within the mutagenesis regions; 2) nonsense variants must match the amber codon UAG; 3) log10 variant frequency is > -5.8; 4) SNP variants with false positive random forest predictions (p > 0.5) in more than half of the gDNA libraries were filtered out. Amino acid positions were adjusted to account for the Flag tag such that positions align with the endogenous membrane-COMT protein isoform. \n\n\nDescription of statistical model for converting counts to scores, including normalization:\n\nFrequencies were normalized to the wild-type count at each position (variant count + 0.5 / wild-type count + 0.5) (Rubin et al, 2017). For each variant, normalized frequencies in the negative control (plasmid template) were subtracted from the frequency observed in total gDNA, total RNA, polysomal RNA, and flow cytometry gDNA samples, and variants with resultant negative values were filtered out. Normalized frequencies were batch-corrected to account for the library preparation experiment using ComBat (Leek et al, 2022) with default parameters. Total gDNA and total RNA frequencies were adjusted with a model matrix, whereas no model matrix was used for polysomal RNA and flow cytometry gDNA libraries, as readouts to be compared were in different batches. The median among technical replicates was computed for each biological replicate (independent stable cell line).\n\nWe only considered variants that were observed in all biological replicates for each readout to be compared such that median log-frequency after wild-type normalization and batch correction was greater than -10. These frequencies were converted to counts per million and input to ALDEx2 (Fernandes et al, 2013, 2014; Gloor et al, 2016; Gloor, 2021) functions aldex.clr, aldex.effect, and aldex.ttest with parameters `paired=FALSE, denom=”iqlr”, mc.samples=128`.\n\nALDEx2 effect size is an estimate of the median standardized difference between groups. To avoid any distributional assumptions for standardization, ALDEx2 uses a permutation based non-parametric estimate of dispersion. We opted to use ALDEx2 75% confidence intervals instead of reporting standard deviations on fold changes. The confidence intervals are determined by Monte Carlo methods that produce a posterior probability distribution of the observed data given repeated sampling. The comparisons made were 1) total RNA to total gDNA; 2) polysome metafraction 3 or 4 to total RNA; 3) flow cytometry population 3 (low moxGFP fluorescence compared to mCherry fluorescence) to total gDNA.\n\n\nDescription of additional data columns included in the score or count tables, including column naming conventions:\n\nscore: ALDEx2 effect size estimate.\nscore_ci_lower: lower 75% confidence interval of ALDEx2 effect size estimate.\nscore_ci_higher: higher 75% confidence interval of ALDEx2 effect size estimate.\nscore_FDR: false discovery rate of ALDEx2 effect size estimate.\n","extraMetadata":{},"recordType":"Experiment","urn":"urn:mavedb:00000661-b","createdBy":{"orcidId":"0000-0003-2307-3299","firstName":"Ian","lastName":"Hoskins","recordType":"User"},"modifiedBy":{"orcidId":"0000-0003-2307-3299","firstName":"Ian","lastName":"Hoskins","recordType":"User"},"creationDate":"2023-11-13","modificationDate":"2023-11-13","publishedDate":"2023-11-13","experimentSetUrn":"urn:mavedb:00000661","doiIdentifiers":[{"identifier":"10.1101/2023.08.02.551517","id":57,"recordType":"DoiIdentifier","url":"https://doi.org/10.1101/2023.08.02.551517"}],"primaryPublicationIdentifiers":[],"secondaryPublicationIdentifiers":[],"rawReadIdentifiers":[{"identifier":"GSE246139","id":87,"recordType":"RawReadIdentifier","url":"http://www.ebi.ac.uk/ena/data/view/GSE246139"}],"contributors":[],"keywords":[],"scoreSetUrns":["urn:mavedb:00000661-b-1"],"externalLinks":{},"numScoreSets":1,"processingState":null,"officialCollections":[]},{"title":"COMT RNA abundance","shortDescription":"Integration of MAVEs to measure post-transcriptional gene expression for >3000 variants in COMT","abstractText":"We integrated three multiplexed assays to measure RNA abundance, protein abundance, and ribosome load for variants in the early coding region of the membrane-COMT (MB-COMT) protein isoform. Two regions that span clinically relevant variants (rs4633, rs4818, rs4680) in MB-COMT were targeted: codons 40-74 (region 1) and 136-158 (region 2). Our design enabled testing the importance of synonymous variation and a potential role of RNA secondary structure in controlling COMT gene expression.","methodText":"Variant library construction methods:\n\nPrimers used to generate the transgene are in Supplementary Table 4 located at https://github.com/ijhoskins/integrated_MAVEs/blob/main/supplementary/Supplementary_Tables.xlsx\n\nA plasmid containing the COMT noncoding exon 2 (NM_000754.4) and coding sequence was obtained from DNASU (clone HsCD00617865), and gBlocks were ordered for 1) the same COMT noncoding exon 2 with a downstream Flag tag; and 2) moxGFP (Costantini et al, 2015), which facilitates membrane localization. The 5′ UTR and coding sequence were amplified in Q5 polymerase reactions (NEB, M0491) with 10 ng HsCD00617865 template and 500 nM forward and reverse primers for 25 cycles using manufacturer’s recommendations with 55 ̊C annealing temperature. Products were purified with the Nucleospin PCR and Gel Cleanup Kit (Takara, 740609), and the 5′ UTR, CDS, and moxGFP gBlock were joined using a Gibson assembly reaction (NEB) following manufacturer recommendations.\n\n2 μL of the assembly reaction was used as template for a NEB Q5 PCR using Gateway-compatible primers with attB1/attB2 tails and the aforementioned PCR conditions. The product was subsequently transferred into a pDONR223 entry clone via a Gateway BP reaction (Thermo Fisher Scientific), and transferred to a destination vector (pDEST_HC_Rec_Bxb_v2) compatible with recombination into the 293T LLP iCasp9 Blast cell line with a 1 h LR reaction (Thermo Fisher Scientific, 11-791-020). Plasmid maps for pDONR223 and pDEST_HC_Rec_Bxb_v2 are located at https://github.com/ijhoskins/integrated_MAVEs/tree/main/plasmid_maps. The isolated clone had a deletion in the 5′ UTR, so a new clone was generated using the deletion clone as a template.\n\nA CDS-moxGFP product was amplified from the deletion clone in a 50 μL NEB Q5 reaction for 25 cycles with 30 s annealing at 55 ̊C. The product was extracted from a 1% agarose-TBE gel (Zymo Gel Extraction kit, 10 μL elution), then assembled with the 5′ UTR gBlock in a Gibson reaction. A final 50 μL Q5 PCR reaction was performed with 2 μL assembly reaction template for 30 cycles with 30 s annealing at 50 ̊C. The product was extracted from a 1% agarose-TBE gel as previously detailed. The transgene was transferred to the pDONR223 and pDEST_HC_Rec_Bxb_v2 with 1 h Gateway BP and LR reactions, respectively. The isolated transgene contained the full 91 nt of exon 2 of the MB-COMT 5′ UTR, a N-terminal Flag tag, a C-terminal moxGFP tag upstream of a bicistronic IRES-mCherry element, and partial vector-derived 5′ and 3′ UTRs from the landing pad vector pDEST_HC_Rec_Bxb_v2.\n\nThe 5′ UTR-Flag-COMT-moxGFP clone in pDEST_HC_Rec_Bxb_v2 was sent to Twist Biosciences for mutagenesis of membrane-COMT codons 40-74 (region 1) and 136-158 (region 2) using degenerate primers. All silent and missense variants and the amber nonsense variant were generated at each codon in the template with added attB recombination sites. Inserts for each target region (region 1, region 2) were received and pooled separately, with codons in each region pooled at equimolar ratios. A Gateway BP reaction was performed with 70 ng of pooled inserts and 150 ng pDONR223 at 25 ̊C overnight (19 h). 1.5 μL BP reaction was electroporated into 25 μL Endura Electrocompetent cells (Lucigen, 60242) with the following conditions: 2 mm cuvette, 2500 V, 200 Ohms, 25 μF. Cells were outgrown for 1 h at 37 ̊C in 1 mL Lucigen outgrowth media, then 600 μL was plated on Nunc Square Bioassay dishes, scraped in 8 mL LB Miller broth, and plasmid library purified with the ZymoPURE II Plasmid Maxiprep Kit (Zymo Research, D4203) using 200 μL 50 ̊C elution buffer, and including endotoxin removal. Library sizes were estimated at 942,000 and 510,000 clones for region 1 and region 2 libraries, respectively (or 434-fold and 357-fold coverage of each variant). 1 μg of each library in pDONR223 was digested with Blp1 (NEB) for 1 h 15 min at 37 ̊C. Digests were run on a 0.8% agarose-TAE gel and full-length plasmids were extracted with the Nucleospin PCR and Gel Cleanup Kit (Takara, 740609). A LR reaction was performed with 150 ng entry library and 150 ng pDEST_HC_Rec_Bxb_v2, incubated at 25 ̊C for 21 h. 25 μL Endura Electrocompetent cells were transformed with 1.5 μL LR reaction using the same conditions as for the BP reaction. After 14.5 h, colonies were collected with 7 mL LB Miller broth, and 3 mL was pelleted and processed with the ZymoPURE II Plasmid Maxiprep Kit.\n\nDescription of functional assay, including model system and selection type:\n\nFor each region 1 and region 2 biological replicate cell line, 20 μg of the library along with an equal mass of Bxb1 recombinase (pCAG-NLS-HA-Bxb1) was transfected into a 15 cm dish of HEK293T LLP iCasp9 Blast cells using Lipofectamine 3000 (Thermo Fisher Scientific, L3000008), with volumes scaled based on 3.75 μL reagent per 6-well. After 48 h, at near full confluency, 2 μg/mL Doxycycline and 10 nM AP1903 (MedChemExpress, HY-16046), both solubilized in DMSO, were added for negative selection of non-recombined cells. The next day, dead cells were removed and recombined cells were grown out for an additional two days with fresh media containing Doxycycline and AP1903. Cells were recovered for two days by growth in media without Doxycycline and AP1903. Before functional assays, transcription was induced with 2 μg/mL Doxycycline for 24 h (total RNA, polysome RNA readouts) or 21 h (flow cytometry readouts), and cells were stimulated with fresh media for 1 h 15 min prior to harvest. Importantly, because the COMT transgene is membrane localized and faces the extracellular space, cells are collected by pipetting in cold PBS, and not by trypsinization.\n\nPolysome fractionation and polysomal RNA purification\nFor COMT library stable cell lines, cells at 80-95% confluency (one 15 cm dish per gradient) were incubated with 100 μg/mL cycloheximide (Sigma Aldrich, C4859) in media for 10 min at 37 ̊C, then resuspended in ice cold PBS with 100 μg/mL cycloheximide (PBS-CHX), collected at 4 ̊C by pipetting, and pelleted at 300 x g for 7 min at 4 ̊C. Pellets were flash frozen in liquid nitrogen, then lysed on ice for 10 min with 300-400 μL lysis buffer (20 mM Tris-HCl pH 7.5; 150 mM NaCl; 5 mM MgCl2, 1 mM fresh DTT, 20 mM ribonucleoside vanadyl complex (Fisher Scientific, 50-812-650), 1X protease inhibitor cocktail EDTA-free (Fisher Scientific, 539196), and 100 μg/ml cycloheximide. Lysates were clarified by centrifugation at 1,300 x g for 10 min at 4 ̊C, and 1/10th of the lysate was saved for total RNA extraction. 20-50% sucrose gradients were made with the BioComp Gradient Maker in 20 mM Tris-HCl pH 7.5, 150 mM NaCl, 5 mM MgCl2, 1 mM DTT, and 100 μg/ml cycloheximide using Beckman Coulter UC tubes 9/16 x 3-1/2 (Fisher Scientific, NC9194790), and equilibrating overnight at 4 ̊C. Clarified lysates were loaded onto sucrose gradients, and fractionated by ultracentrifugation at 38,000 rpm for 2.5 h at 4 ̊C in a Beckman Coulter SW-41 Ti rotor. Approximately 500 μL polysome fractions were collected using the Biocomp Piston Gradient Fractionator (v8.04) with a Triax flow cell (BioComp model FC-1) and Gilson Fraction Collector using a piston speed of 0.2 mm/s. RNA was precipitated from raw polysome fractions by addition of sodium acetate to 300 mM pH 5.2, 2 μL Glycoblue, and 2x volumes cold ethanol, stored overnight at -20  ̊C. RNA was resuspended in 50-80 μL ultrapure water, and these raw fractions were pooled into metafractions and extracted by phenol-chloroform (5:1) and ethanol precipitation overnight at -20 ̊C. After a single 70% ethanol wash, polysomal RNA was resuspended in 30 μL ultrapure water. Lastly, DNaseI treatment and a column purification was performed prior to library preparation.\n\nFlow cytometry analysis and flow cytometry sorting\nFollowing induction with Doxycycline and media stimulation, cells were washed once with cold PBS, resuspended with PBS, and assayed in a 5 mL Falcon polystyrene round bottom tube with a cell-strainer cap (Fisher Scientific, 08-771-23). Analysis of point mutants and deletion cell lines was conducted on a LSRFortessa SORP instrument (BD) with FACSDiva v6.1.3, and flow cytometry was conducted on a FACSAria Fusion SORP instrument (BD) at the Center for Biomedical Research Support Microscopy and Imaging Facility at UT Austin (RRID# SCR_021756). Gating was performed with FlowJo. A gate for cells versus debris was defined based on FSC-H versus SSC-H, and another gate for single cells versus doublets was defined based on FSC-A versus FSC-H. Fluorescence gating was based on autofluorescence of 239T LLP iCasp9 Blast cells without recombination or of a stable cell line expressing the template transgene but without Doxycycline induction. After flow cytometry sorting, each population was expanded up to a confluent T-25 flask, washed, trypsinized, pelleted, and flash frozen prior to gDNA extraction.\n\n\nSequencing strategy and sequencing technology:\n\ngDNA extraction\ngDNA was extracted from approximately 3-4 million cells with the Cell and Tissue DNA Isolation Kit (Norgen Biotek Corp, 24700), including RNaseA treatment at 37 ̊C for 15 min, and eluting in 200 μL warm elution buffer.\n\nRNA extraction and cDNA synthesis\nApproximately 3-4 million cells were solubilized with 1 mL QIAzol (Qiagen, 79306) and 0.2 mL chloroform in 5PRIME Phase-Lock Gel heavy tubes (QuantaBio, 2302830), according to the manufacturer’s recommendations. RNA was precipitated at -20 ̊C following the addition of 2 μL GlycoBlue (Thermo Fisher, AM9515) and 2.5 volumes of cold absolute ethanol. RNA was washed once with cold 70% ethanol then resuspended in 30 μL water. 10 μg total RNA or 15 μL of polysomal RNA (half of 30 μL eluate, volumetric inputs) was treated with DNaseI (NEB, M0303) in a 100 μL reaction at 37 ̊C for 15 min, then re-purified by the RNA Clean and Concentrator Kit (Zymo Research, R1015) and eluted in 20 μL water. 10 μL DNaseI-treated total RNA (<10 μg) or DNaseI-treated polysomal RNA (200 ng - 5 μg) was denatured at 65 ̊C for 5 min followed by RT primer annealing at 4 ̊C for 2 min, using 2 pmol pDEST_HC_Rec_Bxb_v2_R primer specific for the landing pad (see Supplementary Table 4) and 2.5 mM random hexamers. Primed total RNA was included in 20 μL SuperScript IV cDNA synthesis reaction (Thermo Fisher, 18090010) with SUPERase-In RNase inhibitor (Thermo Fisher, AM2696), and first-strand cDNA was synthesized by incubating at 23 ̊C for 10 min, 55 ̊C for 1 h, followed by RT inactivation at 80 ̊C for 10 min. RNA was digested with addition of 5 U RNaseH (NEB, M0297) to the first strand cDNA synthesis reaction and incubation at 37 ̊C for 20 min.\n \nLanding pad amplification (PCR1) \nNucleic acids (gDNA, 1st strand cDNA) were amplified with Q5 polymerase (NEB, M0491) in 50 μL PCR reactions for 16 cycles with 500 nM landing-pad-specific primers (pDEST_HC_Rec_Bxb_v2_F/R) flanking the entire COMT insert (~1.7 kb). For negative control (plasmid) samples, 10 ng was used as template. For total RNA and polysome RNA samples, 15 μL of unpurified 1st strand cDNA synthesis reaction was used as template. For total gDNA samples, 4 μg template was used in each of three PCR reactions. For flow-sorted gDNA samples, either 1.75 μg template was used in three PCR reactions or 2.5 μg was input into two reactions, with 14 cycles instead of 16. The cycling parameters were: initial denaturation at 98 ̊C for 30 s; 3-step cycling with denaturation at 98 ̊C for 10 s, anneal at 65 ̊C for 30 s, extension at 72 ̊C for 1 min; final extension at 72 ̊C for 2 min. Products from total RNA and polysome RNA samples were purified with the QIAquick PCR Purification Kit (Qiagen, 28106) and eluted in 30 μL elution buffer. Products for total gDNA samples were pooled, resolved on a 0.8% agarose/TAE gel, stained with 1x SYBR Gold, and extracted using the Nucleospin Gel and PCR Cleanup Kit (Takara, 740609) or Zymoclean Gel DNA Recovery Kit (Zymo Research, D4007) with 20 μL 70 ̊C elution buffer. At this point, PCR1 products for gDNA samples cannot typically be seen in the gel, and the expected band size is cut using comparison to a high-molecular weight ladder.\n\nCoding sequence amplification (PCR2) \nPurified PCR1 products (landing pad insert) were amplified for each of COMT target regions in a 50 μL NEB Q5 reaction (NEB, M0491) for 8 cycles, following the same cycling parameters as for PCR1 and 500 nM tailed primers (see Supplementary Table 4). 15 μL PCR1 product was used as template (50% of eluate for total RNA and polysome RNA samples; 75% of eluate for gel-extracted gDNA samples). Products were purified with the Nucleospin Gel and PCR Cleanup Kit (Takara, 740609) with 1:5 buffer NTI dilution and 25 μL 70 ̊C elution buffer, or with the QIAquick PCR Purification Kit (Qiagen, 28106) with 30 μL elution buffer.\n\nIllumina adapter addition (PCR3) \nA final Q5 PCR reaction was carried out for 8 cycles with the same formulation as PCR2 but using NEBNext Multiplex Oligos for Illumina Dual Index Primers Set 1 (NEB, E7600S) according to manufacturer’s recommendations (65 ̊C annealing). 25 μL/30 μL eluate was used as template for total gDNA, total RNA, and polysome RNA samples. For flow cytometry gDNA samples, 15 μL/25 μL eluate was used as template. PCR3 products (final library) were purified with the Nucleospin Gel and PCR Cleanup kit (Takara, 740609) and eluted in 25 μL 70 ̊C buffer.\n\nQuantification and size-selection of final libraries \nFinal libraries were quantified using the KAPA Library Quantification Kit for Illumina (Roche, KK4873) according to the manufacturer’s recommendations, using 1:10,000 dilution of libraries and size-correction with an average fragment size of 300 nt (region 1) and 285 nt (region 2). qPCR was performed on the Applied Biosystems ViiA 7 instrument and values were used to pool individual libraries at equimolar ratios to target ~5 M read pairs per library. Final library pools were size-selected use PAGE purification on a 15% TBE gel followed by a crush-and-soak method (Sambrook & Russell, 2006), or with BluePippin 2% gel (Sage Science, BDF2010) to select fragments from 250-340 bp. Following size-selection, final pools were quantified by High Sensitivity DNA Kit (Agilent, 5067-4626) prior to sequencing.\n\nNext-Generation Sequencing\nAll libraries were sequenced 2 x 150 bp (paired-end) on Illumina platforms with minimum 5% PhiX. Libraries were sequenced at MedGenome, Inc. on the HiSeq X or at Novogene Corporation, Inc. on the NovaSeq 6000. Raw FASTQs were obtained directly from the vendor.\n\nStructure of biological and technical replicates:\nFour biological replicate cell lines for each target ROI were generated, with some biological replicates having technical replication (replication of growth, library preparation, and sequencing). The median variant frequency of technical replicates was taken prior to statistical testing with ALDEx2 as described below. \n\n\nSequencing read filtering approach:\n\nVariants were called with satmut_utils v1.0.3-dev001 (Hoskins, 2023). Analysis of COMT libraries prepared with the amplicon method utilized the satmut_utils call parameters `-n 2 -m 1 --r1_fiveprime_adapters TACACGACGCTCTTCCGATCT --r1_threeprime_adapters AGATCGGAAGAGCACACGTCT --r2_fiveprime_adapters AGACGTGTGCTCTTCCGATCT --r2_threeprime_adapters AGATCGGAAGAGCGTCGTGTA`.\n\nRandom forest models were trained on simulated data as previously described for libraries prepared from the amplicon method (Hoskins et al, 2023). Variants were post-processed by the following criteria: variants are within the mutagenesis regions; 2) nonsense variants must match the amber codon UAG; 3) log10 variant frequency is > -5.8; 4) SNP variants with false positive random forest predictions (p > 0.5) in more than half of the gDNA libraries were filtered out. Amino acid positions were adjusted to account for the Flag tag such that positions align with the endogenous membrane-COMT protein isoform. \n\n\nDescription of statistical model for converting counts to scores, including normalization:\n\nFrequencies were normalized to the wild-type count at each position (variant count + 0.5 / wild-type count + 0.5) (Rubin et al, 2017). For each variant, normalized frequencies in the negative control (plasmid template) were subtracted from the frequency observed in total gDNA, total RNA, polysomal RNA, and flow cytometry gDNA samples, and variants with resultant negative values were filtered out. Normalized frequencies were batch-corrected to account for the library preparation experiment using ComBat (Leek et al, 2022) with default parameters. Total gDNA and total RNA frequencies were adjusted with a model matrix, whereas no model matrix was used for polysomal RNA and flow cytometry gDNA libraries, as readouts to be compared were in different batches. The median among technical replicates was computed for each biological replicate (independent stable cell line).\n\nWe only considered variants that were observed in all biological replicates for each readout to be compared such that median log-frequency after wild-type normalization and batch correction was greater than -10. These frequencies were converted to counts per million and input to ALDEx2 (Fernandes et al, 2013, 2014; Gloor et al, 2016; Gloor, 2021) functions aldex.clr, aldex.effect, and aldex.ttest with parameters `paired=FALSE, denom=”iqlr”, mc.samples=128`.\n\nALDEx2 effect size is an estimate of the median standardized difference between groups. To avoid any distributional assumptions for standardization, ALDEx2 uses a permutation based non-parametric estimate of dispersion. We opted to use ALDEx2 75% confidence intervals instead of reporting standard deviations on fold changes. The confidence intervals are determined by Monte Carlo methods that produce a posterior probability distribution of the observed data given repeated sampling. The comparisons made were 1) total RNA to total gDNA; 2) polysome metafraction 3 or 4 to total RNA; 3) flow cytometry population 3 (low moxGFP fluorescence compared to mCherry fluorescence) to total gDNA.\n\n\nDescription of additional data columns included in the score or count tables, including column naming conventions:\n\nscore: ALDEx2 effect size estimate.\nscore_ci_lower: lower 75% confidence interval of ALDEx2 effect size estimate.\nscore_ci_higher: higher 75% confidence interval of ALDEx2 effect size estimate.\nscore_FDR: false discovery rate of ALDEx2 effect size estimate.\n","extraMetadata":{},"recordType":"Experiment","urn":"urn:mavedb:00000661-c","createdBy":{"orcidId":"0000-0003-2307-3299","firstName":"Ian","lastName":"Hoskins","recordType":"User"},"modifiedBy":{"orcidId":"0000-0003-2307-3299","firstName":"Ian","lastName":"Hoskins","recordType":"User"},"creationDate":"2023-11-13","modificationDate":"2023-11-13","publishedDate":"2023-11-13","experimentSetUrn":"urn:mavedb:00000661","doiIdentifiers":[{"identifier":"10.1101/2023.08.02.551517","id":57,"recordType":"DoiIdentifier","url":"https://doi.org/10.1101/2023.08.02.551517"}],"primaryPublicationIdentifiers":[],"secondaryPublicationIdentifiers":[],"rawReadIdentifiers":[{"identifier":"GSE246139","id":87,"recordType":"RawReadIdentifier","url":"http://www.ebi.ac.uk/ena/data/view/GSE246139"}],"contributors":[],"keywords":[],"scoreSetUrns":["urn:mavedb:00000661-c-1"],"externalLinks":{},"numScoreSets":1,"processingState":null,"officialCollections":[]},{"title":"COMT protein abundance","shortDescription":"Integration of MAVEs to measure post-transcriptional gene expression for >3000 variants in COMT","abstractText":"We integrated three multiplexed assays to measure RNA abundance, protein abundance, and ribosome load for variants in the early coding region of the membrane-COMT (MB-COMT) protein isoform. Two regions that span clinically relevant variants (rs4633, rs4818, rs4680) in MB-COMT were targeted: codons 40-74 (region 1) and 136-158 (region 2). Our design enabled testing the importance of synonymous variation and a potential role of RNA secondary structure in controlling COMT gene expression.","methodText":"Variant library construction methods:\n\nPrimers used to generate the transgene are in Supplementary Table 4 located at https://github.com/ijhoskins/integrated_MAVEs/blob/main/supplementary/Supplementary_Tables.xlsx\n\nA plasmid containing the COMT noncoding exon 2 (NM_000754.4) and coding sequence was obtained from DNASU (clone HsCD00617865), and gBlocks were ordered for 1) the same COMT noncoding exon 2 with a downstream Flag tag; and 2) moxGFP (Costantini et al, 2015), which facilitates membrane localization. The 5′ UTR and coding sequence were amplified in Q5 polymerase reactions (NEB, M0491) with 10 ng HsCD00617865 template and 500 nM forward and reverse primers for 25 cycles using manufacturer’s recommendations with 55 ̊C annealing temperature. Products were purified with the Nucleospin PCR and Gel Cleanup Kit (Takara, 740609), and the 5′ UTR, CDS, and moxGFP gBlock were joined using a Gibson assembly reaction (NEB) following manufacturer recommendations.\n\n2 μL of the assembly reaction was used as template for a NEB Q5 PCR using Gateway-compatible primers with attB1/attB2 tails and the aforementioned PCR conditions. The product was subsequently transferred into a pDONR223 entry clone via a Gateway BP reaction (Thermo Fisher Scientific), and transferred to a destination vector (pDEST_HC_Rec_Bxb_v2) compatible with recombination into the 293T LLP iCasp9 Blast cell line with a 1 h LR reaction (Thermo Fisher Scientific, 11-791-020). Plasmid maps for pDONR223 and pDEST_HC_Rec_Bxb_v2 are located at https://github.com/ijhoskins/integrated_MAVEs/tree/main/plasmid_maps. The isolated clone had a deletion in the 5′ UTR, so a new clone was generated using the deletion clone as a template.\n\nA CDS-moxGFP product was amplified from the deletion clone in a 50 μL NEB Q5 reaction for 25 cycles with 30 s annealing at 55 ̊C. The product was extracted from a 1% agarose-TBE gel (Zymo Gel Extraction kit, 10 μL elution), then assembled with the 5′ UTR gBlock in a Gibson reaction. A final 50 μL Q5 PCR reaction was performed with 2 μL assembly reaction template for 30 cycles with 30 s annealing at 50 ̊C. The product was extracted from a 1% agarose-TBE gel as previously detailed. The transgene was transferred to the pDONR223 and pDEST_HC_Rec_Bxb_v2 with 1 h Gateway BP and LR reactions, respectively. The isolated transgene contained the full 91 nt of exon 2 of the MB-COMT 5′ UTR, a N-terminal Flag tag, a C-terminal moxGFP tag upstream of a bicistronic IRES-mCherry element, and partial vector-derived 5′ and 3′ UTRs from the landing pad vector pDEST_HC_Rec_Bxb_v2.\n\nThe 5′ UTR-Flag-COMT-moxGFP clone in pDEST_HC_Rec_Bxb_v2 was sent to Twist Biosciences for mutagenesis of membrane-COMT codons 40-74 (region 1) and 136-158 (region 2) using degenerate primers. All silent and missense variants and the amber nonsense variant were generated at each codon in the template with added attB recombination sites. Inserts for each target region (region 1, region 2) were received and pooled separately, with codons in each region pooled at equimolar ratios. A Gateway BP reaction was performed with 70 ng of pooled inserts and 150 ng pDONR223 at 25 ̊C overnight (19 h). 1.5 μL BP reaction was electroporated into 25 μL Endura Electrocompetent cells (Lucigen, 60242) with the following conditions: 2 mm cuvette, 2500 V, 200 Ohms, 25 μF. Cells were outgrown for 1 h at 37 ̊C in 1 mL Lucigen outgrowth media, then 600 μL was plated on Nunc Square Bioassay dishes, scraped in 8 mL LB Miller broth, and plasmid library purified with the ZymoPURE II Plasmid Maxiprep Kit (Zymo Research, D4203) using 200 μL 50 ̊C elution buffer, and including endotoxin removal. Library sizes were estimated at 942,000 and 510,000 clones for region 1 and region 2 libraries, respectively (or 434-fold and 357-fold coverage of each variant). 1 μg of each library in pDONR223 was digested with Blp1 (NEB) for 1 h 15 min at 37 ̊C. Digests were run on a 0.8% agarose-TAE gel and full-length plasmids were extracted with the Nucleospin PCR and Gel Cleanup Kit (Takara, 740609). A LR reaction was performed with 150 ng entry library and 150 ng pDEST_HC_Rec_Bxb_v2, incubated at 25 ̊C for 21 h. 25 μL Endura Electrocompetent cells were transformed with 1.5 μL LR reaction using the same conditions as for the BP reaction. After 14.5 h, colonies were collected with 7 mL LB Miller broth, and 3 mL was pelleted and processed with the ZymoPURE II Plasmid Maxiprep Kit.\n\nDescription of functional assay, including model system and selection type:\n\nFor each region 1 and region 2 biological replicate cell line, 20 μg of the library along with an equal mass of Bxb1 recombinase (pCAG-NLS-HA-Bxb1) was transfected into a 15 cm dish of HEK293T LLP iCasp9 Blast cells using Lipofectamine 3000 (Thermo Fisher Scientific, L3000008), with volumes scaled based on 3.75 μL reagent per 6-well. After 48 h, at near full confluency, 2 μg/mL Doxycycline and 10 nM AP1903 (MedChemExpress, HY-16046), both solubilized in DMSO, were added for negative selection of non-recombined cells. The next day, dead cells were removed and recombined cells were grown out for an additional two days with fresh media containing Doxycycline and AP1903. Cells were recovered for two days by growth in media without Doxycycline and AP1903. Before functional assays, transcription was induced with 2 μg/mL Doxycycline for 24 h (total RNA, polysome RNA readouts) or 21 h (flow cytometry readouts), and cells were stimulated with fresh media for 1 h 15 min prior to harvest. Importantly, because the COMT transgene is membrane localized and faces the extracellular space, cells are collected by pipetting in cold PBS, and not by trypsinization.\n\nPolysome fractionation and polysomal RNA purification\nFor COMT library stable cell lines, cells at 80-95% confluency (one 15 cm dish per gradient) were incubated with 100 μg/mL cycloheximide (Sigma Aldrich, C4859) in media for 10 min at 37 ̊C, then resuspended in ice cold PBS with 100 μg/mL cycloheximide (PBS-CHX), collected at 4 ̊C by pipetting, and pelleted at 300 x g for 7 min at 4 ̊C. Pellets were flash frozen in liquid nitrogen, then lysed on ice for 10 min with 300-400 μL lysis buffer (20 mM Tris-HCl pH 7.5; 150 mM NaCl; 5 mM MgCl2, 1 mM fresh DTT, 20 mM ribonucleoside vanadyl complex (Fisher Scientific, 50-812-650), 1X protease inhibitor cocktail EDTA-free (Fisher Scientific, 539196), and 100 μg/ml cycloheximide. Lysates were clarified by centrifugation at 1,300 x g for 10 min at 4 ̊C, and 1/10th of the lysate was saved for total RNA extraction. 20-50% sucrose gradients were made with the BioComp Gradient Maker in 20 mM Tris-HCl pH 7.5, 150 mM NaCl, 5 mM MgCl2, 1 mM DTT, and 100 μg/ml cycloheximide using Beckman Coulter UC tubes 9/16 x 3-1/2 (Fisher Scientific, NC9194790), and equilibrating overnight at 4 ̊C. Clarified lysates were loaded onto sucrose gradients, and fractionated by ultracentrifugation at 38,000 rpm for 2.5 h at 4 ̊C in a Beckman Coulter SW-41 Ti rotor. Approximately 500 μL polysome fractions were collected using the Biocomp Piston Gradient Fractionator (v8.04) with a Triax flow cell (BioComp model FC-1) and Gilson Fraction Collector using a piston speed of 0.2 mm/s. RNA was precipitated from raw polysome fractions by addition of sodium acetate to 300 mM pH 5.2, 2 μL Glycoblue, and 2x volumes cold ethanol, stored overnight at -20  ̊C. RNA was resuspended in 50-80 μL ultrapure water, and these raw fractions were pooled into metafractions and extracted by phenol-chloroform (5:1) and ethanol precipitation overnight at -20 ̊C. After a single 70% ethanol wash, polysomal RNA was resuspended in 30 μL ultrapure water. Lastly, DNaseI treatment and a column purification was performed prior to library preparation.\n\nFlow cytometry analysis and flow cytometry sorting\nFollowing induction with Doxycycline and media stimulation, cells were washed once with cold PBS, resuspended with PBS, and assayed in a 5 mL Falcon polystyrene round bottom tube with a cell-strainer cap (Fisher Scientific, 08-771-23). Analysis of point mutants and deletion cell lines was conducted on a LSRFortessa SORP instrument (BD) with FACSDiva v6.1.3, and flow cytometry was conducted on a FACSAria Fusion SORP instrument (BD) at the Center for Biomedical Research Support Microscopy and Imaging Facility at UT Austin (RRID# SCR_021756). Gating was performed with FlowJo. A gate for cells versus debris was defined based on FSC-H versus SSC-H, and another gate for single cells versus doublets was defined based on FSC-A versus FSC-H. Fluorescence gating was based on autofluorescence of 239T LLP iCasp9 Blast cells without recombination or of a stable cell line expressing the template transgene but without Doxycycline induction. After flow cytometry sorting, each population was expanded up to a confluent T-25 flask, washed, trypsinized, pelleted, and flash frozen prior to gDNA extraction.\n\n\nSequencing strategy and sequencing technology:\n\ngDNA extraction\ngDNA was extracted from approximately 3-4 million cells with the Cell and Tissue DNA Isolation Kit (Norgen Biotek Corp, 24700), including RNaseA treatment at 37 ̊C for 15 min, and eluting in 200 μL warm elution buffer.\n\nRNA extraction and cDNA synthesis\nApproximately 3-4 million cells were solubilized with 1 mL QIAzol (Qiagen, 79306) and 0.2 mL chloroform in 5PRIME Phase-Lock Gel heavy tubes (QuantaBio, 2302830), according to the manufacturer’s recommendations. RNA was precipitated at -20 ̊C following the addition of 2 μL GlycoBlue (Thermo Fisher, AM9515) and 2.5 volumes of cold absolute ethanol. RNA was washed once with cold 70% ethanol then resuspended in 30 μL water. 10 μg total RNA or 15 μL of polysomal RNA (half of 30 μL eluate, volumetric inputs) was treated with DNaseI (NEB, M0303) in a 100 μL reaction at 37 ̊C for 15 min, then re-purified by the RNA Clean and Concentrator Kit (Zymo Research, R1015) and eluted in 20 μL water. 10 μL DNaseI-treated total RNA (<10 μg) or DNaseI-treated polysomal RNA (200 ng - 5 μg) was denatured at 65 ̊C for 5 min followed by RT primer annealing at 4 ̊C for 2 min, using 2 pmol pDEST_HC_Rec_Bxb_v2_R primer specific for the landing pad (see Supplementary Table 4) and 2.5 mM random hexamers. Primed total RNA was included in 20 μL SuperScript IV cDNA synthesis reaction (Thermo Fisher, 18090010) with SUPERase-In RNase inhibitor (Thermo Fisher, AM2696), and first-strand cDNA was synthesized by incubating at 23 ̊C for 10 min, 55 ̊C for 1 h, followed by RT inactivation at 80 ̊C for 10 min. RNA was digested with addition of 5 U RNaseH (NEB, M0297) to the first strand cDNA synthesis reaction and incubation at 37 ̊C for 20 min.\n \nLanding pad amplification (PCR1) \nNucleic acids (gDNA, 1st strand cDNA) were amplified with Q5 polymerase (NEB, M0491) in 50 μL PCR reactions for 16 cycles with 500 nM landing-pad-specific primers (pDEST_HC_Rec_Bxb_v2_F/R) flanking the entire COMT insert (~1.7 kb). For negative control (plasmid) samples, 10 ng was used as template. For total RNA and polysome RNA samples, 15 μL of unpurified 1st strand cDNA synthesis reaction was used as template. For total gDNA samples, 4 μg template was used in each of three PCR reactions. For flow-sorted gDNA samples, either 1.75 μg template was used in three PCR reactions or 2.5 μg was input into two reactions, with 14 cycles instead of 16. The cycling parameters were: initial denaturation at 98 ̊C for 30 s; 3-step cycling with denaturation at 98 ̊C for 10 s, anneal at 65 ̊C for 30 s, extension at 72 ̊C for 1 min; final extension at 72 ̊C for 2 min. Products from total RNA and polysome RNA samples were purified with the QIAquick PCR Purification Kit (Qiagen, 28106) and eluted in 30 μL elution buffer. Products for total gDNA samples were pooled, resolved on a 0.8% agarose/TAE gel, stained with 1x SYBR Gold, and extracted using the Nucleospin Gel and PCR Cleanup Kit (Takara, 740609) or Zymoclean Gel DNA Recovery Kit (Zymo Research, D4007) with 20 μL 70 ̊C elution buffer. At this point, PCR1 products for gDNA samples cannot typically be seen in the gel, and the expected band size is cut using comparison to a high-molecular weight ladder.\n\nCoding sequence amplification (PCR2) \nPurified PCR1 products (landing pad insert) were amplified for each of COMT target regions in a 50 μL NEB Q5 reaction (NEB, M0491) for 8 cycles, following the same cycling parameters as for PCR1 and 500 nM tailed primers (see Supplementary Table 4). 15 μL PCR1 product was used as template (50% of eluate for total RNA and polysome RNA samples; 75% of eluate for gel-extracted gDNA samples). Products were purified with the Nucleospin Gel and PCR Cleanup Kit (Takara, 740609) with 1:5 buffer NTI dilution and 25 μL 70 ̊C elution buffer, or with the QIAquick PCR Purification Kit (Qiagen, 28106) with 30 μL elution buffer.\n\nIllumina adapter addition (PCR3) \nA final Q5 PCR reaction was carried out for 8 cycles with the same formulation as PCR2 but using NEBNext Multiplex Oligos for Illumina Dual Index Primers Set 1 (NEB, E7600S) according to manufacturer’s recommendations (65 ̊C annealing). 25 μL/30 μL eluate was used as template for total gDNA, total RNA, and polysome RNA samples. For flow cytometry gDNA samples, 15 μL/25 μL eluate was used as template. PCR3 products (final library) were purified with the Nucleospin Gel and PCR Cleanup kit (Takara, 740609) and eluted in 25 μL 70 ̊C buffer.\n\nQuantification and size-selection of final libraries \nFinal libraries were quantified using the KAPA Library Quantification Kit for Illumina (Roche, KK4873) according to the manufacturer’s recommendations, using 1:10,000 dilution of libraries and size-correction with an average fragment size of 300 nt (region 1) and 285 nt (region 2). qPCR was performed on the Applied Biosystems ViiA 7 instrument and values were used to pool individual libraries at equimolar ratios to target ~5 M read pairs per library. Final library pools were size-selected use PAGE purification on a 15% TBE gel followed by a crush-and-soak method (Sambrook & Russell, 2006), or with BluePippin 2% gel (Sage Science, BDF2010) to select fragments from 250-340 bp. Following size-selection, final pools were quantified by High Sensitivity DNA Kit (Agilent, 5067-4626) prior to sequencing.\n\nNext-Generation Sequencing\nAll libraries were sequenced 2 x 150 bp (paired-end) on Illumina platforms with minimum 5% PhiX. Libraries were sequenced at MedGenome, Inc. on the HiSeq X or at Novogene Corporation, Inc. on the NovaSeq 6000. Raw FASTQs were obtained directly from the vendor.\n\nStructure of biological and technical replicates:\nFour biological replicate cell lines for each target ROI were generated, with some biological replicates having technical replication (replication of growth, library preparation, and sequencing). The median variant frequency of technical replicates was taken prior to statistical testing with ALDEx2 as described below. \n\n\nSequencing read filtering approach:\n\nVariants were called with satmut_utils v1.0.3-dev001 (Hoskins, 2023). Analysis of COMT libraries prepared with the amplicon method utilized the satmut_utils call parameters `-n 2 -m 1 --r1_fiveprime_adapters TACACGACGCTCTTCCGATCT --r1_threeprime_adapters AGATCGGAAGAGCACACGTCT --r2_fiveprime_adapters AGACGTGTGCTCTTCCGATCT --r2_threeprime_adapters AGATCGGAAGAGCGTCGTGTA`.\n\nRandom forest models were trained on simulated data as previously described for libraries prepared from the amplicon method (Hoskins et al, 2023). Variants were post-processed by the following criteria: variants are within the mutagenesis regions; 2) nonsense variants must match the amber codon UAG; 3) log10 variant frequency is > -5.8; 4) SNP variants with false positive random forest predictions (p > 0.5) in more than half of the gDNA libraries were filtered out. Amino acid positions were adjusted to account for the Flag tag such that positions align with the endogenous membrane-COMT protein isoform. \n\n\nDescription of statistical model for converting counts to scores, including normalization:\n\nFrequencies were normalized to the wild-type count at each position (variant count + 0.5 / wild-type count + 0.5) (Rubin et al, 2017). For each variant, normalized frequencies in the negative control (plasmid template) were subtracted from the frequency observed in total gDNA, total RNA, polysomal RNA, and flow cytometry gDNA samples, and variants with resultant negative values were filtered out. Normalized frequencies were batch-corrected to account for the library preparation experiment using ComBat (Leek et al, 2022) with default parameters. Total gDNA and total RNA frequencies were adjusted with a model matrix, whereas no model matrix was used for polysomal RNA and flow cytometry gDNA libraries, as readouts to be compared were in different batches. The median among technical replicates was computed for each biological replicate (independent stable cell line).\n\nWe only considered variants that were observed in all biological replicates for each readout to be compared such that median log-frequency after wild-type normalization and batch correction was greater than -10. These frequencies were converted to counts per million and input to ALDEx2 (Fernandes et al, 2013, 2014; Gloor et al, 2016; Gloor, 2021) functions aldex.clr, aldex.effect, and aldex.ttest with parameters `paired=FALSE, denom=”iqlr”, mc.samples=128`.\n\nALDEx2 effect size is an estimate of the median standardized difference between groups. To avoid any distributional assumptions for standardization, ALDEx2 uses a permutation based non-parametric estimate of dispersion. We opted to use ALDEx2 75% confidence intervals instead of reporting standard deviations on fold changes. The confidence intervals are determined by Monte Carlo methods that produce a posterior probability distribution of the observed data given repeated sampling. The comparisons made were 1) total RNA to total gDNA; 2) polysome metafraction 3 or 4 to total RNA; 3) flow cytometry population 3 (low moxGFP fluorescence compared to mCherry fluorescence) to total gDNA.\n\n\nDescription of additional data columns included in the score or count tables, including column naming conventions:\n\nscore: ALDEx2 effect size estimate.\nscore_ci_lower: lower 75% confidence interval of ALDEx2 effect size estimate.\nscore_ci_higher: higher 75% confidence interval of ALDEx2 effect size estimate.\nscore_FDR: false discovery rate of ALDEx2 effect size estimate.\n","extraMetadata":{},"recordType":"Experiment","urn":"urn:mavedb:00000661-d","createdBy":{"orcidId":"0000-0003-2307-3299","firstName":"Ian","lastName":"Hoskins","recordType":"User"},"modifiedBy":{"orcidId":"0000-0003-2307-3299","firstName":"Ian","lastName":"Hoskins","recordType":"User"},"creationDate":"2023-11-13","modificationDate":"2023-11-13","publishedDate":"2023-11-13","experimentSetUrn":"urn:mavedb:00000661","doiIdentifiers":[{"identifier":"10.1101/2023.08.02.551517","id":57,"recordType":"DoiIdentifier","url":"https://doi.org/10.1101/2023.08.02.551517"}],"primaryPublicationIdentifiers":[],"secondaryPublicationIdentifiers":[],"rawReadIdentifiers":[{"identifier":"GSE246139","id":87,"recordType":"RawReadIdentifier","url":"http://www.ebi.ac.uk/ena/data/view/GSE246139"}],"contributors":[],"keywords":[],"scoreSetUrns":["urn:mavedb:00000661-d-1"],"externalLinks":{},"numScoreSets":1,"processingState":null,"officialCollections":[]}],"createdBy":{"orcidId":"0000-0003-2307-3299","firstName":"Ian","lastName":"Hoskins","recordType":"User"},"modifiedBy":{"orcidId":"0000-0003-2307-3299","firstName":"Ian","lastName":"Hoskins","recordType":"User"},"creationDate":"2023-11-13","modificationDate":"2023-11-13","contributors":[],"numExperiments":4}