diff --git a/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software.md b/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software.md index 60c9a82b0..ed3c142e7 100644 --- a/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software.md +++ b/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software.md @@ -3,20 +3,19 @@ | 3rd Party Software Name | License | License URL | Copyright Notice | | :---------------------- | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :--------------------------------------------------------------------------------------------------------------------------------------------------------- | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | R | [GNU GENERAL PUBLIC LICENSE Version 2, June 1991, and Version 3, 29 June 2007](Microarray_Affymetrix_3rd_Party_Software_Licenses/R_GPL-2_and_GPL-3_LICENSES.pdf) | [https://www.r-project.org/Licenses/](https://www.r-project.org/Licenses/) | Version 2: Copyright (C) 1989, 1991 Free Software Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA Everyone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed.; Version 3: Copyright (C) 2007 Free Software Foundation, Inc. http://fsf.org/ Everyone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed. | -| DT | [GNU GENERAL PUBLIC LICENSE Version 3, 29 June 2007](Microarray_Affymetrix_3rd_Party_Software_Licenses/DT_LICENSE.pdf) | [https://github.com/rstudio/DT/blob/main/LICENSE](https://github.com/rstudio/DT/blob/main/LICENSE) | Version 3: Copyright (C) 2007 Free Software Foundation, Inc. http://fsf.org/ Everyone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed. | -| dplyr | [MIT License Copyright (c) 2022 dplyr authors](Microarray_Affymetrix_3rd_Party_Software_Licenses/dplyr_LICENSE.pdf) | [https://github.com/tidyverse/dplyr/blob/main/LICENSE.md](https://github.com/tidyverse/dplyr/blob/main/LICENSE.md) | Copyright (c) 2022 dplyr authors Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: | -| stringr | [MIT License Copyright (c) 2020 stringr authors](Microarray_Affymetrix_3rd_Party_Software_Licenses/stringr_LICENSE.pdf) | [https://github.com/tidyverse/stringr/blob/main/LICENSE.md](https://github.com/tidyverse/stringr/blob/main/LICENSE.md) | Copyright (c) 2020 stringr authors Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: | -| R.utils | [GNU LESSER GENERAL PUBLIC LICENSE Version 2.1, February 1999](Microarray_Affymetrix_3rd_Party_Software_Licenses/R-utils_LICENSE.pdf) | [https://github.com/HenrikBengtsson/R.utils/blob/2.12.2/DESCRIPTION](https://github.com/HenrikBengtsson/R.utils/blob/2.12.2/DESCRIPTION) | Copyright (C) 1991, 1999 Free Software Foundation, Inc. 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA Everyone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed. | -| oligo | [GNU LESSER GENERAL PUBLIC LICENSE Version 2.1, February 1999](Microarray_Affymetrix_3rd_Party_Software_Licenses/oligo_LICENSE.pdf) | [https://www.bioconductor.org/packages/3.14/bioc/html/oligo.html](https://www.bioconductor.org/packages/3.14/bioc/html/oligo.html) | Copyright (C) 1991, 1999 Free Software Foundation, Inc. 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA Everyone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed. | -| limma | [GNU GENERAL PUBLIC LICENSE Version 2, June 1991](Microarray_Affymetrix_3rd_Party_Software_Licenses/limma_LICENSE.pdf) | [https://bioconductor.org/packages/3.14/bioc/html/limma.html](https://bioconductor.org/packages/3.14/bioc/html/limma.html) | Version 2: Copyright (C) 1989, 1991 Free Software Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA Everyone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed. | -| glue | [MIT License Copyright (c) 2021 glue authors](Microarray_Affymetrix_3rd_Party_Software_Licenses/glue_LICENSE.pdf) | [https://github.com/tidyverse/glue/blob/v1.6.2/LICENSE.md](https://github.com/tidyverse/glue/blob/v1.6.2/LICENSE.md) | Copyright (c) 2021 glue authors Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: | -| biomaRt | [The Artistic License 2.0 ](Microarray_Affymetrix_3rd_Party_Software_Licenses/biomaRt_LICENSE.pdf) | [https://bioconductor.org/packages/3.14/bioc/manuals/biomaRt/man/biomaRt.pdf](https://bioconductor.org/packages/3.14/bioc/manuals/biomaRt/man/biomaRt.pdf) | Copyright (c) 2000-2006, The Perl FoundationEveryone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed. | -| matrixStats | [The Artistic License 2.0 ](Microarray_Affymetrix_3rd_Party_Software_Licenses/matrixStats_LICENSE.pdf) | [https://github.com/HenrikBengtsson/matrixStats/blob/develop/DESCRIPTION](https://github.com/HenrikBengtsson/matrixStats/blob/develop/DESCRIPTION) | Copyright (c) 2000-2006, The Perl FoundationEveryone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed. | -| statmod | [GNU GENERAL PUBLIC LICENSE Version 3, 29 June 2007](Microarray_Affymetrix_3rd_Party_Software_Licenses/statmod_LICENSE.pdf) | [https://github.com/cran/statmod/blob/1.5.0/DESCRIPTION](https://github.com/cran/statmod/blob/1.5.0/DESCRIPTION) | Version 3: Copyright (C) 2007 Free Software Foundation, Inc. http://fsf.org/ Everyone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed. | +| DT | [MIT LICENSE Copyright (c) 2025 DT authors](Microarray_Affymetrix_3rd_Party_Software_Licenses/DT_LICENSE.pdf) | [https://github.com/rstudio/DT/blob/v0.34.0/LICENSE.md](https://github.com/rstudio/DT/blob/v0.34.0/LICENSE.md) | Copyright (c) 2025 DT authors Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: | +| dplyr | [MIT License Copyright (c) 2026 dplyr authors](Microarray_Affymetrix_3rd_Party_Software_Licenses/dplyr_LICENSE.pdf) | [https://github.com/tidyverse/dplyr/blob/v1.2.0/LICENSE.md](https://github.com/tidyverse/dplyr/blob/v1.2.0/LICENSE.md) | Copyright (c) 2026 dplyr authors Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: | +| stringr | [MIT License Copyright (c) 2023 stringr authors](Microarray_Affymetrix_3rd_Party_Software_Licenses/stringr_LICENSE.pdf) | [https://github.com/tidyverse/stringr/blob/v1.6.0/LICENSE.md](https://github.com/tidyverse/stringr/blob/v1.6.0/LICENSE.md) | Copyright (c) 2023 stringr authors Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: | +| oligo | [GNU LESSER GENERAL PUBLIC LICENSE Version 2.1, February 1999](Microarray_Affymetrix_3rd_Party_Software_Licenses/oligo_LICENSE.pdf) | [https://www.bioconductor.org/packages/3.22/bioc/html/oligo.html](https://www.bioconductor.org/packages/3.22/bioc/html/oligo.html) | Copyright (C) 1991, 1999 Free Software Foundation, Inc. 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA Everyone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed. | +| limma | [GNU GENERAL PUBLIC LICENSE Version 2, June 1991](Microarray_Affymetrix_3rd_Party_Software_Licenses/limma_LICENSE.pdf) | [https://bioconductor.org/packages/3.22/bioc/html/limma.html](https://bioconductor.org/packages/3.22/bioc/html/limma.html) | Version 2: Copyright (C) 1989, 1991 Free Software Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA Everyone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed. | +| glue | [MIT License Copyright (c) 2023 glue authors](Microarray_Affymetrix_3rd_Party_Software_Licenses/glue_LICENSE.pdf) | [https://github.com/tidyverse/glue/blob/v1.8.0/LICENSE.md](https://github.com/tidyverse/glue/blob/v1.8.0/LICENSE.md) | Copyright (c) 2023 glue authors Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: | +| purrr | [MIT License Copyright (c) 2023 purrr authors](Microarray_Affymetrix_3rd_Party_Software_Licenses/purrr_LICENSE.pdf) | [https://github.com/tidyverse/purrr/blob/v1.2.1/LICENSE.md](https://github.com/tidyverse/purrr/blob/v1.2.1/LICENSE.md) | Copyright (c) 2023 purrr authors Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: | +| matrixStats | [The Artistic License 2.0 ](Microarray_Affymetrix_3rd_Party_Software_Licenses/matrixStats_LICENSE.pdf) | [https://github.com/HenrikBengtsson/matrixStats/blob/1.5.0/DESCRIPTION](https://github.com/HenrikBengtsson/matrixStats/blob/1.5.0/DESCRIPTION) | Copyright (c) 2000-2006, The Perl FoundationEveryone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed. | +| statmod | [GNU GENERAL PUBLIC LICENSE Version 3, 29 June 2007](Microarray_Affymetrix_3rd_Party_Software_Licenses/statmod_LICENSE.pdf) | [https://github.com/cran/statmod/blob/1.5.1/DESCRIPTION](https://github.com/cran/statmod/blob/1.5.1/DESCRIPTION) | Version 3: Copyright (C) 2007 Free Software Foundation, Inc. http://fsf.org/ Everyone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed. | | Singularity | [The BSD 3-clause License (MIT) Copyright (c) 2015-2017, Gregory M. Kurtzer. Copyright (c) 2016-2017, The Regents of the University of California. Copyright (c) 2017, SingularityWare, LLC. Copyright (c) 2018-2022, Sylabs, Inc.](Microarray_Affymetrix_3rd_Party_Software_Licenses/Singularity_LICENSE.pdf) | [https://docs.sylabs.io/guides/3.9/user-guide/license.html](https://docs.sylabs.io/guides/3.9/user-guide/license.html) | Copyright (c) 2015-2017, Gregory M. Kurtzer. Copyright (c) 2016-2017, The Regents of the University of California. Copyright (c) 2017, SingularityWare, LLC. Copyright (c) 2018-2022, Sylabs, Inc. Redistribution and use in source and binary forms, with or without modification, are permitted provided that the following conditions are met: | | dp_tools | [The MIT License (MIT) Copyright (c) 2022 Jonathan Oribello](Microarray_Affymetrix_3rd_Party_Software_Licenses/dp_tools_LICENSE.pdf) | [https://github.com/J-81/dp_tools/blob/main/LICENSE](https://github.com/J-81/dp_tools/blob/main/LICENSE) | Copyright (c) 2022 Jonathan Oribello. Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the “Software”), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: | -| Quarto | [GNU GENERAL PUBLIC LICENSE Version 2, June 1991](Microarray_Affymetrix_3rd_Party_Software_Licenses/Quarto_LICENSE.pdf) | [https://quarto.org/license.html](https://quarto.org/license.html) | Version 2: Copyright (C) 1989, 1991 Free Software Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA Everyone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed. | -| tibble | [MIT License Copyright (c) 2019 RStudio and others](Microarray_Affymetrix_3rd_Party_Software_Licenses/tibble_LICENSE.pdf) | [https://tibble.tidyverse.org/LICENSE.html](https://tibble.tidyverse.org/LICENSE.html) | Copyright (c) 2019 RStudio and others. Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the “Software”), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions | +| Quarto | [MIT License Copyright (c) 2020-2024 Posit Software, PBC authors](Microarray_Affymetrix_3rd_Party_Software_Licenses/Quarto_LICENSE.pdf) | [https://github.com/quarto-dev/quarto-cli/blob/v1.9.36/COPYING.md](https://github.com/quarto-dev/quarto-cli/blob/v1.9.36/COPYING.md) | Copyright (c) 2020-2024 Posit Software, PBC authors Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: | +| tibble | [MIT License Copyright (c) 2019 RStudio and others](Microarray_Affymetrix_3rd_Party_Software_Licenses/tibble_LICENSE.pdf) | [https://github.com/tidyverse/tibble/blob/v3.2.1/LICENSE.md](https://github.com/tidyverse/tibble/blob/v3.2.1/LICENSE.md) | Copyright (c) 2019 RStudio and others. Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the “Software”), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions | |Nextflow | [Apache License Version 2.0, January 2004](Microarray_Affymetrix_3rd_Party_Software_Licenses/Nextflow_LICENSE.pdf) | [https://github.com/nextflow-io/nextflow/blob/master/COPYING](https://github.com/nextflow-io/nextflow/blob/master/COPYING) | Grant of Copyright License. Subject to the terms and conditions of this License, each Contributor hereby grants to You a perpetual, worldwide, non-exclusive, no-charge, royalty-free, irrevocable copyright license to reproduce, prepare Derivative Works of, publicly display, publicly perform, sublicense, and distribute the Work and such Derivative Works in Source or Object form. | |python | [The PSF LICENSE AGREEMENT Copyright (c) 2001-2022 Python Software Foundation](Microarray_Affymetrix_3rd_Party_Software_Licenses/PYTHON_LICENSE.pdf) | [https://docs.python.org/3/license.html#psf-license-agreement-for-python-release](https://docs.python.org/3/license.html#psf-license-agreement-for-python-release) | Copyright (c) 2001-2022 Python Software Foundation; All Rights Reserved | |pandas | [The BSD 3-clause License](Microarray_Affymetrix_3rd_Party_Software_Licenses/PANDAS_LICENSE.pdf)| [https://github.com/pandas-dev/pandas/blob/main/LICENSE](https://github.com/pandas-dev/pandas/blob/main/LICENSE)| Copyright (c) 2008-2011, AQR Capital Management, LLC, Lambda Foundry, Inc. and PyData Development Team; All rights reserved. Copyright (c) 2011-2022, Open source contributors. Redistribution and use in source and binary forms, with or without modification, are permitted provided that the following conditions are met:| diff --git a/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/DT_LICENSE.pdf b/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/DT_LICENSE.pdf index eee3cb2a3..b64ce24fb 100644 Binary files a/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/DT_LICENSE.pdf and b/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/DT_LICENSE.pdf differ diff --git a/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/Quarto_LICENSE.pdf b/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/Quarto_LICENSE.pdf index dd8b58044..362d5c3a1 100644 Binary files a/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/Quarto_LICENSE.pdf and b/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/Quarto_LICENSE.pdf differ diff --git a/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/R-utils_LICENSE.pdf b/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/R-utils_LICENSE.pdf deleted file mode 100644 index e48248c35..000000000 Binary files a/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/R-utils_LICENSE.pdf and /dev/null differ diff --git a/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/biomaRt_LICENSE.pdf b/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/biomaRt_LICENSE.pdf deleted file mode 100644 index 7b45d07a4..000000000 Binary files a/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/biomaRt_LICENSE.pdf and /dev/null differ diff --git a/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/dplyr_LICENSE.pdf b/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/dplyr_LICENSE.pdf index cc0416a3b..905eabde1 100644 Binary files a/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/dplyr_LICENSE.pdf and b/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/dplyr_LICENSE.pdf differ diff --git a/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/glue_LICENSE.pdf b/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/glue_LICENSE.pdf index 2e1518ed9..f492380e1 100644 Binary files a/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/glue_LICENSE.pdf and b/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/glue_LICENSE.pdf differ diff --git a/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/purrr_LICENSE.pdf b/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/purrr_LICENSE.pdf new file mode 100644 index 000000000..4b7c7a6fa Binary files /dev/null and b/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/purrr_LICENSE.pdf differ diff --git a/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/stringr_LICENSE.pdf b/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/stringr_LICENSE.pdf index 0e97d1804..40afc74f6 100644 Binary files a/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/stringr_LICENSE.pdf and b/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/stringr_LICENSE.pdf differ diff --git a/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/tibble_LICENSE.pdf b/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/tibble_LICENSE.pdf index 1d3231d9b..eeb03aa03 100644 Binary files a/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/tibble_LICENSE.pdf and b/3rd_Party_Licenses/Microarray_Affymetrix_3rd_Party_Software_Licenses/tibble_LICENSE.pdf differ diff --git a/Microarray/Affymetrix/Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114-A.md b/Microarray/Affymetrix/Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114-A.md new file mode 100644 index 000000000..a0c424c59 --- /dev/null +++ b/Microarray/Affymetrix/Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114-A.md @@ -0,0 +1,1540 @@ +# GeneLab bioinformatics processing pipeline for Affymetrix microarray data + +> **This page holds an overview and instructions for how GeneLab processes Affymetrix microarray datasets. Exact processing commands and GL-DPPD-7114 version used for specific GeneLab datasets (GLDS) are provided with their processed data in the [Open Science Data Repository (OSDR)](https://osdr.nasa.gov/bio/repo).** +> +> \* The pipeline detailed below is currently used for animal and *Arabidopsis thaliana* studies only, it will be updated soon for processing microbe microarray data and other plant data. + +--- + +**Date:** May XX, 2026 +**Revision:** -A +**Document Number:** GL-DPPD-7114-A + +**Submitted by:** +Crystal Han and Jihan Yehia (GeneLab Data Processing Team) + +**Approved by:** +Jonathan Galazka (OSDR Project Manager) +Danielle Lopez (OSDR Deputy Project Manager) +Amanda Saravia-Butler (OSDR Subject Matter Expert) +Barbara Novak (GeneLab Data Processing Lead) + +--- + +## Updates from previous version + +Updated [Ensembl Reference Files](https://github.com/nasa/GeneLab_Data_Processing/blob/master/GeneLab_Reference_Annotations/Pipeline_GL-DPPD-7110_Versions/GL-DPPD-7110-A/GL-DPPD-7110-A_annotations.csv) to the following releases: +- Animals: Ensembl release 112 +- Plants: Ensembl plants release 59 +- Bacteria: Ensembl bacteria release 59 + +Software Updates: + +| Program | Previous Version | New Version | +| :----------- | :--------------- | :---------- | +| R | 4.1.3 | 4.5.3 | +| DT | 0.26 | 0.34.0 | +| dplyr | 1.0.10 | 1.2.0 | +| glue | 1.6.2 | 1.8.0 | +| tibble | 3.1.8 | 3.3.1 | +| stringr | 1.5.0 | 1.6.0 | +| purrr | 1.0.1 | 1.2.1 | +| Bioconductor | 3.14 | 3.22 | +| oligo | 1.58.0 | 1.74.0 | +| limma | 3.50.3 | 3.66.0 | +| matrixStats | 0.63.0 | 1.5.0 | +| statmod | 1.5.0 | 1.5.1 | +| dp_tools | 1.3.4 | 1.3.8 | +| Quarto | 1.2.313 | 1.9.36 | + +- Packages `R.utils` and `biomaRt` were removed from the pipeline as they are no longer used in the processing code. + +Code Changes: + +- Added support for plotting HTAFeatureSet data in MA plots, see [Step 3c](#3c-ma-plots) + +- Added ability to use custom gene annotations when annotations are not available in Ensembl FTP, see [Step 8](#8-probeset-annotations) + +- Replaced biomaRt's live Ensembl query with a static Ensembl FTP download for probe-to-gene mapping in [Step 8](#8a-probeset-annotations), to avoid issues with biomaRt's live Ensembl query being down or unavailable + - Gene/transcript ID column position is not stable across organisms in the Ensembl FTP mart dump. Columns are therefore detected by content pattern (matching unversioned Ensembl stable IDs) rather than hardcoded position, making this function organism-agnostic. + +- Simplified group sample retrieval in [Step 9c](#9c-add-annotation-and-stats-columns-and-format-de-table) to use a more concise `filter/pull/sort` chain instead of `group_by/summarize/filter/pull`, addressing the deprecation warning in dplyr >= 1.1.0 where returning more than 1 row per `summarise()` group is deprecated + +--- + +# Table of contents + +- [Software used](#software-used) +- [General processing overview with example commands](#general-processing-overview-with-example-commands) + - [1. Create Sample RunSheet](#1-create-sample-runsheet) + - [2. Load Data](#2-load-data) + - [2a. Load Libraries and Define Input Parameters](#2a-load-libraries-and-define-input-parameters) + - [2b. Define Custom Functions](#2b-define-custom-functions) + - [2c. Load Metadata and Raw Data](#2c-load-metadata-and-raw-data) + - [3. Raw Data Quality Assessment](#3-raw-data-quality-assessment) + - [3a. Density Plot](#3a-density-plot) + - [3b. Pseudo Image Plots](#3b-pseudo-image-plots) + - [3c. MA Plots](#3c-ma-plots) + - [3d. Boxplots](#3d-boxplots) + - [4. Background Correction](#4-background-correction) + - [5. Between Array Normalization](#5-between-array-normalization) + - [6. Normalized Data Quality Assessment](#6-normalized-data-quality-assessment) + - [6a. Density Plot](#6a-density-plot) + - [6b. Pseudo Image Plots](#6b-pseudo-image-plots) + - [6c. MA Plots](#6c-ma-plots) + - [6d. Boxplots](#6d-boxplots) + - [7. Probeset Summarization](#7-probeset-summarization) + - [8. Probeset Annotations](#8-probeset-annotations) + - [8a. Get Probeset Annotations](#8a-get-probeset-annotations) + - [8b. Summarize Gene Mapping](#8b-summarize-gene-mapping) + - [8c. Generate Annotated Raw and Normalized Expression Tables](#8c-generate-annotated-raw-and-normalized-expression-tables) + - [9. Probeset Differential Expression (DE)](#9-probeset-differential-expression-de) + - [9a. Generate Design Matrix](#9a-generate-design-matrix) + - [9b. Perform Individual Probeset Level DE](#9b-perform-individual-probeset-level-de) + - [9c. Add Annotation and Stats Columns and Format DE Table](#9c-add-annotation-and-stats-columns-and-format-de-table) + +--- + +# Software used + +| Program | Version | Relevant Links | +| :----------- | :-----: | :------------------------------------------------------------------------------------------------------------------------- | +| R | 4.5.3 | [https://www.r-project.org/](https://www.r-project.org/) | +| DT | 0.34.0 | [https://github.com/rstudio/DT](https://github.com/rstudio/DT) | +| dplyr | 1.2.0 | [https://dplyr.tidyverse.org](https://dplyr.tidyverse.org) | +| glue | 1.8.0 | [https://glue.tidyverse.org](https://glue.tidyverse.org) | +| tibble | 3.3.1 | [https://tibble.tidyverse.org](https://tibble.tidyverse.org) | +| stringr | 1.6.0 | [https://stringr.tidyverse.org](https://stringr.tidyverse.org) | +| purrr | 1.2.1 | [https://purrr.tidyverse.org](https://purrr.tidyverse.org) | +| Bioconductor | 3.22 | [https://bioconductor.org](https://bioconductor.org) | +| oligo | 1.74.0 | [https://bioconductor.org/packages/3.22/bioc/html/oligo.html](https://bioconductor.org/packages/3.22/bioc/html/oligo.html) | +| limma | 3.66.0 | [https://bioconductor.org/packages/3.22/bioc/html/limma.html](https://bioconductor.org/packages/3.22/bioc/html/limma.html) | +| matrixStats | 1.5.0 | [https://github.com/HenrikBengtsson/matrixStats](https://github.com/HenrikBengtsson/matrixStats) | +| statmod | 1.5.1 | [https://github.com/cran/statmod](https://github.com/cran/statmod) | +| dp_tools | 1.3.8 | [https://github.com/torres-alexis/dp_tools](https://github.com/torres-alexis/dp_tools) | +| Quarto | 1.9.36 | [https://quarto.org](https://quarto.org) | +--- + +# General processing overview with example commands + +> [!IMPORTANT] +> Exact processing commands and output files listed in **bold** below are included with each Microarray processed dataset in the [Open Science Data Repository (OSDR)](https://osdr.nasa.gov/bio/repo/). + +--- + +## 1. Create Sample RunSheet + +> [!TIP] +> - Rather than running the commands below to create the runsheet needed for processing, the runsheet may also be created manually by following the [file specification](../Workflow_Documentation/NF_MAAffymetrix/examples/runsheet/README.md). +> +> - These command line tools are part of the [dp_tools](https://github.com/torres-alexis/dp_tools) program. + +```bash +### Download the *ISA.zip file from the GeneLab Repository ### + +dpt-get-isa-archive \ + --accession OSD-### + +### Parse the metadata from the *ISA.zip file to create a sample runsheet ### + +dpt-isa-to-runsheet --accession OSD-### \ + --plugin-dir /path/to/dp_tools__microarray_plugin \ + --isa-archive *ISA.zip +``` + +**Parameter Definitions:** + +- `--accession OSD-###` - OSD accession ID (replace ### with the OSD number being processed), used to retrieve the urls for the ISA archive and raw expression files hosted on the GeneLab Repository +- `--plugin-dir /path/to/dp_tools__microarray_plugin` - specifies the path to the plugin directory defining the dp-tools configuration for the desired assay type. A plugin for both the Affymetrix microarray assay is provided in the [Workflow_Documentation](../Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix/) folder +- `--isa-archive` - Specifies the *ISA.zip file for the respective OSD dataset, downloaded in the `dpt-get-isa-archive` command + + +**Input Data:** + +- No input data required but the OSD accession ID needs to be indicated, which is used to download the respective ISA archive + +**Output Data:** + +- *ISA.zip (compressed ISA directory containing Investigation, Study, and Assay (ISA) metadata files for the respective OSD dataset, used to define sample groups - the *ISA.zip file is located in the [OSD repository](https://osdr.nasa.gov/bio/repo/search?q=&data_source=cgene,alsda&data_type=study) under 'Study Files' -> 'metadata') + +- **{OSD-Accession-ID}_microarray_v{version}_runsheet.csv** (table containing metadata required for processing, version denotes the dp_tools schema used to specify the metadata to extract from the ISA archive) + +
+ +--- + +## 2. Load Data + +> [!NOTE] +> Steps 2 - 9 are done in R + +
+ +### 2a. Load Libraries and Define Input Parameters + +```R +### Install R packages if not already installed ### + +install.packages("dplyr") +install.packages("tibble") +install.packages("stringr") +install.packages("purrr") +install.packages("glue") +install.packages("matrixStats") +install.packages("statmod") +if (!require("BiocManager", quietly = TRUE)) + install.packages("BiocManager") +BiocManager::install(version = "3.22") +BiocManager::install("limma") +BiocManager::install("oligo") + + +## Note: Only dplyr is explicitly loaded. Other library functions are called with explicit namespace (e.g. LIBRARYNAME::FUNCTION) +library(dplyr) # Ensure infix operator is available, methods should still reference dplyr namespace otherwise +options(dplyr.summarise.inform = FALSE) # Don't print out message informing how the result will be grouped + +# Define path to runsheet +runsheet <- "/path/to/runsheet/{OSD-Accession-ID}_microarray_v{version}_runsheet.csv" + +## Set up output structure + +# Output Constants +DIR_RAW_DATA <- "00-RawData" +DIR_NORMALIZED_EXPRESSION <- "01-oligo_NormExp" +DIR_DGE <- "02-limma_DGE" + +dir.create(DIR_RAW_DATA) +dir.create(DIR_NORMALIZED_EXPRESSION) +dir.create(DIR_DGE) + +## Save original par settings +## Par may be temporarily changed for plotting purposes and reset once the plotting is done + +original_par <- par() +options(preferRaster=TRUE) # use Raster when possible to avoid antialiasing artifacts in images + +options(timeout=1000) # ensure enough time for data downloads +``` + +
+ +### 2b. Define Custom Functions + +#### retry_with_delay() +
+ utility function to improve robustness of function calls; used to remedy intermittent internet issues during runtime + + ```R + retry_with_delay <- function(func, ...) { + max_attempts = 5 + initial_delay = 10 + delay_increase = 30 + attempt <- 1 + current_delay <- initial_delay + while (attempt <= max_attempts) { + result <- tryCatch( + expr = func(...), + error = function(e) e + ) + + if (!inherits(result, "error")) { + return(result) + } else { + if (attempt < max_attempts) { + message(paste("Retry attempt", attempt, "failed for function with name <", deparse(substitute(func)) ,">. Retrying in", current_delay, "second(s)...")) + Sys.sleep(current_delay) + current_delay <- current_delay + delay_increase + } else { + stop(paste("Max retry attempts reached. Last error:", result$message)) + } + } + + attempt <- attempt + 1 + } + } + ``` + + **Function Parameter Definitions:** + - `func=` - specifies the function to wrap + - `...` - other arguments passed on to the `func` + + **Returns:** the output of the wrapped function +
+ +#### shortened_organism_name() +
+ shortens organism names, for example 'Homo Sapiens' to 'hsapiens' + + ```R + shortened_organism_name <- function(long_name) { + tokens <- long_name %>% stringr::str_split(" ", simplify = TRUE) + genus_name <- tokens[1] + + species_name <- tokens[2] + + short_name <- stringr::str_to_lower(paste0(substr(genus_name, start = 1, stop = 1), species_name)) + + return(short_name) + } + ``` + + **Function Parameter Definitions:** + - `long_name=` - a string containing the long name of the organism + + **Returns:** a string containing the short name of the organism +
+ +#### get_biomart_attribute() +
+ retrieves resolved BioMart attribute source from runsheet dataframe + + ```R + get_biomart_attribute <- function(df_rs) { + # check if runsheet has 'biomart_attribute' column + if ( !is.null(df_rs$`biomart_attribute`) ) { + print("Using attribute name sourced from runsheet") + # Format according to biomart needs + formatted_value <- unique(df_rs$`biomart_attribute`) %>% + stringr::str_replace_all(" ","_") %>% # Replace all spaces with underscore + stringr::str_to_lower() # Lower casing only + return(formatted_value) + } else { + stop("ERROR: Could not find 'biomart_attribute' in runsheet") + } + } + ``` + + **Function Parameter Definitions:** + - `df_rs=` - a dataframe containing the sample runsheet information + + **Returns:** a string containing the formatted value from the `biomart_attribute` column of the runsheet, with all spaces converted to underscores and uppercase converted to lowercase; if no `biomart_attribute` exists in the runsheet, stop and return an error +
+ +#### resolve_mart_ftp_base() +
+ resolves the FTP mart directory and per-species dataset prefix for either the main Ensembl release or an Ensembl Genomes division + + ```R + resolve_mart_ftp_base <- function(division, organism, ensembl_version, ensembl_genomes_portal = NULL) { + if (division == "genomes") { + list( + ftp_dir = glue::glue("https://ftp.ebi.ac.uk/ensemblgenomes/pub/{ensembl_genomes_portal}/release-{ensembl_version}/mysql/{ensembl_genomes_portal}_mart_{ensembl_version}"), + dataset_prefix = glue::glue("{organism}_eg_gene") + ) + } else { + list( + ftp_dir = glue::glue("https://ftp.ebi.ac.uk/pub/ensembl/release-{ensembl_version}/mysql/ensembl_mart_{ensembl_version}"), + dataset_prefix = glue::glue("{organism}_gene_ensembl") + ) + } + } + ``` + + **Function Parameter Definitions:** + - `division=` - a string containing the Ensembl division, either 'ensembl' for the main Ensembl release or 'genomes' for an Ensembl Genomes division + - `organism=` - a string containing the name of the organism (formatted using `shortened_organism_name()`) + - `ensembl_version=` - a string containing the version of Ensembl to use + - `ensembl_genomes_portal=` - a string containing the name of the genomes portal, for example 'plants'; required only when division = "genomes" + + **Returns:** a list containing the resolved `ftp_dir` (the FTP directory URL for this release/division) and `dataset_prefix` (the per-organism dataset filename prefix) +
+ +#### download_mart_dump() +
+ downloads and loads a single gzipped, tab-separated, headerless Ensembl "mart dump" table + + ```R + download_mart_dump <- function(url, col_names = NULL) { + options(timeout = 300) # Can be further increased for downloading large files + print(glue::glue("Mart dump URL: {url}")) + + # Download the file to a temporary location and read it in + temp_file <- tempfile(fileext = ".gz") + download.file(url = url, destfile = temp_file, method = "libcurl") # libcurl needed for ftp(s) URLs + + # Read the gzipped file into a dataframe, optionally assigning column names to the leading columns + mapping <- read.table(gzfile(temp_file), header = FALSE, sep = "\t", quote = "", comment.char = "") + if (!is.null(col_names)) colnames(mapping)[seq_along(col_names)] <- col_names + unlink(temp_file) + return(mapping) + } + ``` + + **Function Parameter Definitions:** + - `url=` - a string containing the URL of the Ensembl "mart dump" file to download + - `col_names=` - an optional character vector of column names to assign to the leading columns of the loaded table + + **Returns:** a dataframe containing the contents of the downloaded "mart dump" table, with column names assigned as specified +
+ +#### get_transcript_to_gene_mapping_from_ftp +
+ obtains a transcript-to-gene mapping table directly from the Ensembl FTP "transcript main" mart dump + + ```R + get_transcript_to_gene_mapping_from_ftp <- function(division, organism, ensembl_version, ensembl_genomes_portal = NULL) { + # Download the transcript main table from Ensembl FTP and extract the columns containing Ensembl gene and transcript IDs + loc <- resolve_mart_ftp_base(division, organism, ensembl_version, ensembl_genomes_portal) + url <- glue::glue("{loc$ftp_dir}/{loc$dataset_prefix}__transcript__main.txt.gz") + raw <- retry_with_delay(download_mart_dump, url) + + # Identify the columns containing Ensembl gene and transcript IDs by matching their content patterns + gene_col_idx <- which(sapply(raw, function(col) any(grepl("^ENS[A-Z]*G[0-9]+$", col)))) + transcript_col_idx <- which(sapply(raw, function(col) any(grepl("^ENS[A-Z]*T[0-9]+$", col)))) + + stopifnot( + "Could not uniquely identify gene ID column in transcript main table" = length(gene_col_idx) == 1, + "Could not uniquely identify transcript ID column in transcript main table" = length(transcript_col_idx) == 1 + ) + + # Extract only the identified gene and transcript ID columns, ensuring uniqueness + raw %>% + dplyr::transmute( + ensembl_transcript_id = .data[[paste0("V", transcript_col_idx)]], + ensembl_gene_id = .data[[paste0("V", gene_col_idx)]] + ) %>% + dplyr::distinct() + } + ``` + + **Function Parameter Definitions:** + - `division=` - a string containing the Ensembl division ("main" for the main Ensembl release; Ensembl Genomes divisions map probe to gene directly and never call this function) + - `organism=` - a string containing the name of the organism (formatted using `shortened_organism_name()`) + - `ensembl_version=` - a string containing the version of Ensembl to use + - `biomart_attribute=` - a string containing the BioMart attribute (formatted using `get_biomart_attribute()`) + - `ensembl_genomes_portal=` - unused for division = "main"; present for signature consistency with resolve_mart_ftp_base() + + **Returns:** a dataframe mapping Ensembl transcript IDs to Ensembl gene IDs, as obtained via FTP +
+ +#### get_probe_to_gene_mapping_from_ftp +
+ obtains a probe-to-gene mapping table directly from Ensembl FTP mart dumps, for either main Ensembl or an Ensembl Genomes division + + ```R + get_probe_to_gene_mapping_from_ftp <- function(division, organism, ensembl_version, biomart_attribute, ensembl_genomes_portal = NULL) { + # Download the probe mapping table from Ensembl FTP and extract the columns containing Ensembl gene IDs and the specified BioMart attribute + loc <- resolve_mart_ftp_base(division, organism, ensembl_version, ensembl_genomes_portal) + dm_url <- glue::glue("{loc$ftp_dir}/{loc$dataset_prefix}__efg_{biomart_attribute}__dm.txt.gz") + + probe_dm <- tryCatch( + if (division == "genomes") { + download_mart_dump(dm_url, col_names = c("MAPID", "ensembl_gene_id", biomart_attribute)) + } else { + download_mart_dump(dm_url, col_names = c("MAPID", "ensembl_transcript_id", biomart_attribute)) + }, + error = function(e) { + message(glue::glue("Probe mapping file not available for attribute '{biomart_attribute}' ({e$message}); falling back to custom annotation")) + NULL + } + ) + + if (is.null(probe_dm)) return(NULL) + + # If the division is "genomes", the probe mapping table already contains Ensembl gene IDs, return it directly + if (division == "genomes") { + return(probe_dm %>% dplyr::select(!!sym(biomart_attribute), ensembl_gene_id)) + } + + # If the division is "main", join the probe mapping table with the transcript-to-gene mapping table to get Ensembl gene IDs + transcript_to_gene <- get_transcript_to_gene_mapping_from_ftp(division, organism, ensembl_version, ensembl_genomes_portal) + + probe_dm %>% + dplyr::left_join(transcript_to_gene, by = "ensembl_transcript_id") %>% + dplyr::select(!!sym(biomart_attribute), ensembl_gene_id) + } + ``` + + **Function Parameter Definitions:** + - `division=` - a string containing the Ensembl division, either 'ensembl' for the main Ensembl release or 'genomes' for an Ensembl Genomes division + - `organism=` - a string containing the name of the organism (formatted using `shortened_organism_name()`) + - `ensembl_version=` - a string containing the version of Ensembl to use + - `biomart_attribute=` - a string containing the BioMart attribute (formatted using `get_biomart_attribute()`) + - `ensembl_genomes_portal=` - a string containing the name of the genomes portal, for example 'plants'; required only when division = "genomes" + + **Returns:** a dataframe mapping the probe attribute column to ensembl_gene_id; NULL if the array design's attribute-specific dm file is not available on Ensembl FTP +
+ +#### list_to_unique_piped_string() +
+ converts character vector into string denoting unique elements separated by '|' characters + + ```R + list_to_unique_piped_string <- function(str_list) { + #! Convert vector of multi-mapped genes to string separated by '|' characters + #! e.g. c("GO1","GO2","GO2","G03") -> "GO1|GO2|GO3" + return(toString(unique(str_list)) %>% stringr::str_replace_all(pattern = stringr::fixed(", "), replacement = "|")) + } + ``` + + **Function Parameter Definitions:** + - `str_list=` - vector of character elements + + **Returns:** a string containing the unique elements from `str_list` concatenated together, separated by '|' characters +
+ +#### runsheet_to_design_matrix() +
+ loads the GeneLab runsheet into a list of dataframes + + ```R + runsheet_to_design_matrix <- function(runsheet_path) { + # Pull all factors for each sample in the study from the runsheet created in Step 1 + df <- read.csv(runsheet, check.names = FALSE) %>% + dplyr::mutate_all(function(x) iconv(x, "latin1", "ASCII", sub="")) # Convert all characters to ascii, when not possible, remove the character # get only Factor Value columns + factors = as.data.frame(df[,grep("Factor.Value", colnames(df), ignore.case=TRUE)]) + colnames(factors) = paste("factor",1:dim(factors)[2], sep= "_") + + # Load metadata from runsheet csv file + compare_csv = data.frame(sample_id = df[,c("Sample Name")], factors) + + # Create data frame containing all samples and respective factors + study <- as.data.frame(compare_csv[,2:dim(compare_csv)[2]]) + colnames(study) <- colnames(compare_csv)[2:dim(compare_csv)[2]] + rownames(study) <- compare_csv[,1] + + # Format groups and indicate the group that each sample belongs to + if (dim(study)[2] >= 2){ + group<-apply(study,1,paste,collapse = " & ") # concatenate multiple factors into one condition per sample + } else{ + group<-study[,1] + } + group_names <- paste0("(",group,")",sep = "") # human readable group names + group <- sub("^BLOCKER_", "", make.names(paste0("BLOCKER_", group))) # group naming compatible with R models, this maintains the default behavior of make.names with the exception that 'X' is never prepended to group names + names(group) <- group_names + + # Format contrasts table, defining pairwise comparisons for all groups + contrast.names <- combn(levels(factor(names(group))),2) # generate matrix of pairwise group combinations for comparison + contrasts <- apply(contrast.names, MARGIN=2, function(col) sub("^BLOCKER_", "", make.names(paste0("BLOCKER_", stringr::str_sub(col, 2, -2))))) + contrast.names <- c(paste(contrast.names[1,],contrast.names[2,],sep = "v"),paste(contrast.names[2,],contrast.names[1,],sep = "v")) # format combinations for output table files names + contrasts <- cbind(contrasts,contrasts[c(2,1),]) + colnames(contrasts) <- contrast.names + sampleTable <- data.frame(condition=factor(group)) + rownames(sampleTable) <- df[,c("Sample Name")] + + condition <- sampleTable[,'condition'] + names_mapping <- as.data.frame(cbind(safe_name = as.character(condition), original_name = group_names)) + + design <- model.matrix(~ 0 + condition) + design_data <- list( matrix = design, mapping = names_mapping, groups = as.data.frame( cbind(sample = df[,c("Sample Name")], group = group_names) ), contrasts = contrasts ) + return(design_data) + } + ``` + + **Function Parameter Definitions:** + - `runsheet_path=` - a string containing the path to the runsheet generated in [Step 1](#1-create-sample-runsheet) + + **Returns:** a list of R objects containing the sample information and metadata + - `design_data$matrix` - a design (or model) matrix describing the conditions in the dataset + - `design_data$mapping` - a dataframe mapping the human-readable group names to the names of the conditions modified for use in R + - `design_data$groups` - a dataframe of group names and contrasts for each sample + - `design_data$contrasts` - a matrix of all pairwise comparisons of the groups +
+ +#### lm_fit_pairwise() +
+ performs all pairwise comparisons using limma::lmFit() + + ```R + lm_fit_pairwise <- function(norm_data, design) { + # Approach based on limma manual section 17.4 (version 3.52.4) + fit <- limma::lmFit(norm_data, design) + + # Create Contrast Model + fit.groups <- colnames(fit$design)[which(fit$assign == 1)] + combos <- combn(fit.groups,2) + contrasts<-c(paste(combos[1,],combos[2,],sep = "-"),paste(combos[2,],combos[1,],sep = "-")) # format combinations for limma:makeContrasts + cont.matrix <- limma::makeContrasts(contrasts=contrasts,levels=design) + contrast.fit <- limma::contrasts.fit(fit, cont.matrix) + + contrast.fit <- limma::eBayes(contrast.fit,trend=TRUE,robust=TRUE) + return(contrast.fit) + } + ``` + + **Function Parameter Definitions:** + - `norm_data=` - an R object containing log-ratios or log-expression values for a series of arrays, with rows corresponding to genes and columns to samples + - `design=` - the design matrix of the microarray experiment, with rows corresponding to samples and columns to coefficients to be estimated + + **Returns:** an R object of class `MArrayLM` +
+ +#### reformat_names() +
+ reformats column names for consistency across DE analyses tables within GeneLab + + ```R + reformat_names <- function(colname, group_name_mapping) { + new_colname <- colname %>% + stringr::str_replace(pattern = "^P.value.adj.condition", replacement = "Adj.p.value_") %>% + stringr::str_replace(pattern = "^P.value.condition", replacement = "P.value_") %>% + stringr::str_replace(pattern = "^Coef.condition", replacement = "Log2fc_") %>% # This is the Log2FC as per: https://rdrr.io/bioc/limma/man/writefit.html + stringr::str_replace(pattern = "^t.condition", replacement = "T.stat_") %>% + stringr::str_replace(pattern = ".condition", replacement = "v") + + # remap to group names before make.names was applied + unique_group_name_mapping <- unique(group_name_mapping) %>% arrange(-nchar(safe_name)) + for ( i in seq(nrow(unique_group_name_mapping)) ) { + safe_name <- unique_group_name_mapping[i,]$safe_name + original_name <- unique_group_name_mapping[i,]$original_name + new_colname <- new_colname %>% stringr::str_replace(pattern = stringr::fixed(safe_name), replacement = original_name) + } + + return(new_colname) + } + ``` + + **Function Parameter Definitions:** + - `colnames=` - a character vector containing the column names to reformat + - `group_name_mapping=` - a dataframe mapping the original human-readable group names to the R modified safe names + + **Returns:** a character vector containing the formatted column names +
+ +#### generate_prefixed_column_order() +
+ creates a vector of column names based on subject and given prefixes; used for both contrasts and groups column name generation + + ```R + generate_prefixed_column_order <- function(subjects, prefixes) { + # Track order of columns + final_order = c() + + # For each contrast + for (subject in subjects) { + # Generate column names for each prefix and append to final_order + for (prefix in prefixes) { + final_order <- append(final_order, glue::glue("{prefix}{subject}")) + } + } + return(final_order) + } + ``` + + **Function Parameter Definitions:** + - `subjects` - a character vector containing subject strings to add prefixes to + - `prefixes` - a character vector of prefixes to add to the beginning of each subject string + + **Returns:** a character vector with all possible combinations of prefix + subject +
+ +
+ +### 2c. Load Metadata and Raw Data + +```R +df_rs <- read.csv(runsheet, check.names = FALSE) %>% + dplyr::mutate_all(function(x) iconv(x, "latin1", "ASCII", sub="")) # Convert all characters to ascii, when not possible, remove the character + +# Determine expected local filename per sample (Array Data File Path is only checked for ".gz", never used to locate files) +local_paths <- ifelse( + stringr::str_detect(df_rs$`Array Data File Path`, "\\.gz$"), + stringr::str_remove(df_rs$`Array Data File Name`, "\\.gz$"), + df_rs$`Array Data File Name` +) + +df_local_paths <- data.frame(`Sample Name` = df_rs$`Sample Name`, `Local Paths` = local_paths, check.names = FALSE) + +# Load raw data into R object +# Retry with delay here to accomodate oligo's automatic loading of annotation packages and occasional internet related failures to load +raw_data <- retry_with_delay( + oligo::read.celfiles, + df_local_paths$`Local Paths`, + sampleNames = df_local_paths$`Sample Name`# Map column names as Sample Names (instead of default filenames) + ) + +# Summarize raw data +print(paste0("Number of Arrays: ", dim(raw_data)[2])) +print(paste0("Number of Probes: ", dim(raw_data)[1])) +``` + +**Custom Functions Used:** + +- [retry_with_delay()](#retry_with_delay) + +**Input Data:** + +- `runsheet` (Path to runsheet, output from [Step 1](#1-create-sample-runsheet)) +- Raw array data files listed in the runsheet's `Array Data File Name` column, decompressed (if applicable) and placed in the current working directory + +**Output Data:** + +- `df_rs` (R dataframe containing information from the runsheet) +- `raw_data` (R object containing raw microarray data) + + > [!NOTE] + > The raw data R object will be used to generate quality assessment (QA) plots in the next step. + +
+ +--- + +## 3. Raw Data Quality Assessment + +
+ +### 3a. Density Plot + +```R +# Plot settings +par( + xpd = TRUE # Ensure legend can extend past plot area +) + +number_of_sets = ceiling(dim(raw_data)[2] / 30) # Set of 30 samples, used to scale plot +scale_factor = 0.2 # Default scale factor + +if (max(nchar(colnames(raw_data@assayData$exprs))) > 35 & number_of_sets > 1) { # Scale more if sample names are long + scale_factor = if_else(number_of_sets == 2, 0.4, 0.25) +} + +oligo::hist(raw_data, + transfo=log2, # Log2 transform raw intensity values + which=c("all"), + nsample=10000, # Number of probes to plot + main = "Density of raw intensities for multiple arrays") +legend("topright", legend = colnames(raw_data@assayData$exprs), + lty = c(1,2,3,4,5), # Seems like oligo::hist cycles through these first five line types + col = oligo::darkColors(n = ncol(raw_data)), # Ensure legend color is in sync with plot + ncol = number_of_sets, # Set number of columns by number of sets + cex = max(0.35, 1 + scale_factor - (number_of_sets*scale_factor)) # Reduce for each column beyond 1 with minimum of 35% + ) + +# Reset par +par(original_par) +``` + +**Input Data:** + +- `raw_data` (raw data R object created in [Step 2c](#2c-load-metadata-and-raw-data) above) + +**Output Data:** + +- Plot containing the density of raw intensities for each array (lack of overlap indicates a need for normalization) + +
+ +### 3b. Pseudo Image Plots + +```R +for ( i in seq_along(1:ncol(raw_data))) { + oligo::image(raw_data[,i], + transfo = log2, + main = colnames(raw_data)[i] + ) +} +``` + +**Input Data:** + +- `raw_data` (raw data R object created in [Step 2c](#2c-load-metadata-and-raw-data) above) + +**Output Data:** + +- Pseudo images of each array before background correction and normalization + +
+ +### 3c. MA Plots + +```R +if (inherits(raw_data, "GeneFeatureSet")) { + print("Raw data is a GeneFeatureSet, using exprs() to access expression values and adding 0.0001 to avoid log(0)") +} else if (inherits(raw_data, "ExpressionSet") || inherits(raw_data, "ExpressionFeatureSet") || inherits(raw_data, "HTAFeatureSet")) { + print(paste0("Raw data is ", class(raw_data), ". Using default approach for this class for MA Plot")) +} + +if (inherits(raw_data, "GeneFeatureSet")) { + MA_plot <- oligo::MAplot( + exprs(raw_data) + 0.0001, + transfo=log2, + ylim=c(-2, 4), + main="" # This function uses 'main' as a suffix to the sample name. Here we want just the sample name, thus here main is an empty string + ) +} else if (inherits(raw_data, "ExpressionSet") || inherits(raw_data, "ExpressionFeatureSet") || inherits(raw_data, "HTAFeatureSet")) { + MA_plot <- oligo::MAplot( + raw_data, + ylim=c(-2, 4), + main="" # This function uses 'main' as a suffix to the sample name. Here we want just the sample name, thus here main is an empty string + ) +} else { + stop(glue::glue("No strategy for MA plots for {class(raw_data)}")) +} +``` + +**Input Data:** + +- `raw_data` (raw data R object created in [Step 2c](#2c-load-metadata-and-raw-data) above) + +**Output Data:** + +- `MA_plot` (M (log ratio of the subject array vs a pseudo-reference, the median of all other arrays) vs. A (average log expression) plot for each array before background correction and normalization) + +
+ + +### 3d. Boxplots + +```R +max_samplename_length <- max(nchar(colnames(raw_data))) +dynamic_lefthand_margin <- max(max_samplename_length * 0.7, 10) +par( + mar = c(8, dynamic_lefthand_margin, 8, 2) + 0.1, # mar is the margin around the plot. c(bottom, left, top, right) + xpd = TRUE + ) +boxplot <- oligo::boxplot(raw_data[, rev(colnames(raw_data))], # Here we reverse column order to ensure descending order for samples in horizontal boxplot + transfo=log2, # Log2 transform raw intensity values + which=c("all"), + nsample=10000, # Number of probes to plot + las = 1, # las specifies the orientation of the axis labels. 1 = always horizontal + ylab="", + xlab="log2 Intensity", + main = "Boxplot of raw intensities \nfor perfect match and mismatch probes", + horizontal = TRUE + ) +title(ylab = "Sample Name", mgp = c(dynamic_lefthand_margin-2, 1, 0)) +# Reset par +par(original_par) +``` + +**Input Data:** + +- `raw_data` (raw data R object created in [Step 2c](#2c-load-metadata-and-raw-data) above) + +**Output Data:** + +- `boxplot` (Boxplot of raw expression data for each array before background correction and normalization) + +
+ +--- + +## 4. Background Correction + +```R +background_corrected_data <- raw_data %>% oligo::backgroundCorrect(method="rma") +``` + +**Input Data:** + +- `raw_data` (raw data R object created in [Step 2c](#2c-load-metadata-and-raw-data) above) + +**Output Data:** + +- `background_corrected_data` (R object containing background-corrected microarray data) + + > [!NOTE] + > Background correction was performed using the oligo `rma` method, specifically "Convolution Background Correction" + +
+ +--- + +## 5. Between Array Normalization + +```R +# Normalize background-corrected data using the quantile method +norm_data <- oligo::normalize(background_corrected_data, + method = "quantile", + target = "core" # Use oligo default: core metaprobeset mappings + ) + +# Summarize background-corrected and normalized data +print(paste0("Number of Arrays: ", dim(norm_data)[2])) +print(paste0("Number of Probes: ", dim(norm_data)[1])) +``` + +**Input Data:** + +- `background_corrected_data` (R object containing background-corrected microarray data created in [Step 4](#4-background-correction) above) + +**Output Data:** + +- `norm_data` (R object containing background-corrected and normalized microarray data) + + > [!NOTE] + > Normalization was performed using the `quantile` method, which forces the entire empirical distribution of all arrays to be identical + +
+ +--- + +## 6. Normalized Data Quality Assessment + +
+ +### 6a. Density Plot + +```R +# Plot settings +par( + xpd = TRUE # Ensure legend can extend past plot area +) + +number_of_sets = ceiling(dim(norm_data)[2] / 30) # Set of 30 samples, used to scale plot + +oligo::hist(norm_data, + transfo=log2, # Log2 transform normalized intensity values + which=c("all"), + nsample=10000, # Number of probes to plot + main = "Density of normalized intensities for multiple arrays") +legend("topright", legend = colnames(norm_data@assayData$exprs), + lty = c(1,2,3,4,5), # Seems like oligo::hist cycles through these first five line types + col = oligo::darkColors(n = ncol(norm_data)), # Ensure legend color is in sync with plot + ncol = number_of_sets, # Set number of columns by number of sets + cex = max(0.35, 1 + scale_factor - (number_of_sets*scale_factor)) # Reduce for each column beyond 1 with minimum of 35% + ) + +# Reset par +par(original_par) +``` + +**Input Data:** + +- `norm_data` (R object containing background-corrected and normalized microarray data created in [Step 5](#5-between-array-normalization) above) + +**Output Data:** + +- Plot containing the density of background-corrected and normalized intensities for each array (near complete overlap is expected after normalization) + +
+ +### 6b. Pseudo Image Plots + +```R +for ( i in seq_along(1:ncol(norm_data))) { + oligo::image(norm_data[,i], + transfo = log2, + main = colnames(norm_data)[i] + ) +} +``` + +**Input Data:** + +- `norm_data` (R object containing background-corrected and normalized microarray data created in [Step 5](#5-between-array-normalization) above) + +**Output Data:** + +- Pseudo images of each array after background correction and normalization + +
+ +### 6c. MA Plots + +```R +MA_plot <- oligo::MAplot( + norm_data, + ylim=c(-2, 4), + main="" # This function uses 'main' as a suffix to the sample name. Here we want just the sample name, thus here main is an empty string +) +``` + +**Input Data:** + +- `norm_data` (R object containing background-corrected and normalized microarray data created in [Step 5](#5-between-array-normalization) above) + +**Output Data:** + +- `MA_plot` (M (log ratio of the subject array vs a pseudo-reference, the median of all other arrays) vs. A (average log expression) plot for each array after background correction and normalization) + +
+ +### 6d. Boxplots + +```R +max_samplename_length <- max(nchar(colnames(norm_data))) +dynamic_lefthand_margin <- max(max_samplename_length * 0.7, 10) +par( + mar = c(8, dynamic_lefthand_margin, 8, 2) + 0.1, # mar is the margin around the plot. c(bottom, left, top, right) + xpd = TRUE + ) +boxplot <- oligo::boxplot(norm_data[, rev(colnames(norm_data))], # Here we reverse column order to ensure descending order for samples in horizontal boxplot + transfo=log2, # Log2 transform normalized intensity values + which=c("all"), + nsample=10000, # Number of probes to plot + las = 1, # las specifies the orientation of the axis labels. 1 = always horizontal + ylab="", + xlab="log2 Intensity", + main = "Boxplot of normalized intensities \nfor perfect match and mismatch probes", + horizontal = TRUE + ) +title(ylab = "Sample Name", mgp = c(dynamic_lefthand_margin-2, 1, 0)) +# Reset par +par(original_par) +``` + +**Input Data:** + +- `norm_data` (R object containing background-corrected and normalized microarray data created in [Step 5](#5-between-array-normalization) above) + +**Output Data:** + +- `boxplot` (Boxplot of expression data for each array after background correction and normalization) + +
+ +--- + +## 7. Probeset Summarization + +```R +probeset_level_data <- oligo::rma(norm_data, + normalize=FALSE, + background=FALSE + ) + +# Summarize background-corrected and normalized data +print("Summarized Probeset Level Data Below") +print(paste0("Number of Arrays: ", dim(probeset_level_data)[2])) +print(paste0("Total Number of Probes Assigned To A Probeset: ", dim(oligo::getProbeInfo(probeset_level_data, target="core")['man_fsetid'])[1])) # man_fsetid means 'Manufacturer Probeset ID'. Ref: https://support.bioconductor.org/p/57191/ +print(paste0("Number of Probesets: ", dim(unique(oligo::getProbeInfo(probeset_level_data, target="core")['man_fsetid']))[1])) # man_fsetid means 'Manufacturer Probeset ID'. Ref: https://support.bioconductor.org/p/57191/ +``` + +**Input Data:** + +- `norm_data` (R object containing background-corrected and normalized microarray data created in [Step 5](#5-between-array-normalization) above) + +**Output Data:** + +- `probeset_level_data` (R object containing probeset level expression values after summarization of normalized probeset level data) + +
+ +--- + +## 8. Probeset Annotations + +
+ +### 8a. Get Probeset Annotations + +```R +# If using custom annotation, local_annotation_dir is path to directory containing annotation file and and array_annot_path is path/url to file containing design information for the array +local_annotation_dir <- NULL # +array_annot_path <- NULL # + +ENSEMBL_VERSION <- ensembl_version +expected_attribute_name <- get_biomart_attribute(df_rs) + +organism <- shortened_organism_name(unique(df_rs$organism)) +annot_key <- ifelse(organism %in% c("athaliana"), 'TAIR', 'ENSEMBL') + +if (organism %in% c("athaliana")) { + ensembl_genomes_portal = "plants" + print(glue::glue("Using Ensembl Genomes FTP to get probe mapping table. Portal: {ensembl_genomes_portal}, version: {ENSEMBL_VERSION}")) + + df_mapping <- get_probe_to_gene_mapping_from_ftp( + division = "genomes", + organism = organism, + ensembl_version = ENSEMBL_VERSION, + biomart_attribute = expected_attribute_name, + ensembl_genomes_portal = ensembl_genomes_portal + ) + + # TAIR IDs in the mapping tables tend to be in the format 'AT1G01010.1' but the raw data has 'AT1G01010' + # So here we remove the '.NNN' from the mapping table where .NNN is any number + df_mapping$ensembl_gene_id <- stringr::str_replace_all(df_mapping$ensembl_gene_id, "\\.\\d+$", "") + + use_custom_annot <- FALSE +} else { + expected_dataset_name <- glue::glue("{organism}_gene_ensembl") + print(glue::glue("Expected dataset name: '{expected_dataset_name}'")) + print(glue::glue("Expected attribute name: '{expected_attribute_name}'")) + print(glue::glue("Searching for Ensembl Version: {ENSEMBL_VERSION}")) + + # Some probe_ids for affy_hta_2_0 may end in .hg.1 instead of .hg (how it is in biomaRt), leading to 0 results returned + if (expected_attribute_name == 'affy_hta_2_0') { + rownames(probeset_level_data) <- stringr::str_replace(rownames(probeset_level_data), '\\.hg\\.1$', '.hg') + } + + probe_ids <- rownames(probeset_level_data) + + print(glue::glue("Using Ensembl biomart to get specific version of mapping table. Ensembl version: {ENSEMBL_VERSION}")) + print(glue::glue("Attempting Ensembl FTP probe mapping table. Version: {ENSEMBL_VERSION}")) + df_mapping <- get_probe_to_gene_mapping_from_ftp( + division = "main", + organism = organism, + ensembl_version = ENSEMBL_VERSION, + biomart_attribute = expected_attribute_name + ) + + if (!is.null(df_mapping)) { + use_custom_annot <- FALSE + # FTP has no server-side filter equivalent to getBM(filters=, values=) + # the full mart-dump file is always downloaded in full, then filtered + # client-side down to just this experiment's probes. + df_mapping <- df_mapping %>% dplyr::filter(!!sym(expected_attribute_name) %in% probe_ids) + } else { + use_custom_annot <- TRUE + } +} + +# At this point, we have df_mapping from either the Ensembl FTP mart dumps (main or Ensembl Genomes) depending on the organism +# If no df_mapping obtained (e.g., organism not supported in biomart), use custom annotations; otherwise, merge in-house annotations to df_mapping + +if (use_custom_annot) { + expected_attribute_name <- 'ProbesetID' + annot_type <- 'NO_CUSTOM_ANNOT' + + if (!is.null(local_annotation_dir) && !is.null(array_annot_path)) { + probe_annot_df <- read.csv(array_annot_path, row.names=1) + if (unique(df_rs$`biomart_attribute`) %in% row.names(probe_annot_df)) { + annot_config <- probe_annot_df[unique(df_rs$`biomart_attribute`), ] + annot_type <- annot_config$annot_type[[1]] + } else { + warning(paste0("No entry for '", unique(df_rs$`biomart_attribute`), "' in provided custom probe annotation file: ", array_annot_path)) + } + } else { + warning("Need to provide both local_annotation_dir and array_annot_path to use custom annotation.") + } + + if (annot_type == '3prime-IVT') { + unique_probe_ids <- read.csv( + file.path(local_annotation_dir, annot_config$annot_filename[[1]]), + skip = 13, header = TRUE, na.strings = c('NA', '---') + )[c('Probe.Set.ID', 'Entrez.Gene', 'Gene.Symbol', 'Gene.Title', 'Ensembl', 'RefSeq.Transcript.ID', 'RefSeq.Protein.ID', 'Gene.Ontology.Biological.Process', 'Gene.Ontology.Cellular.Component', 'Gene.Ontology.Molecular.Function')] + + # Clean columns + unique_probe_ids$Gene.Symbol <- purrr::map_chr(stringr::str_split(unique_probe_ids$Gene.Symbol, stringr::fixed(' /// ')), ~paste0(unique(.), collapse = "|")) %>% stringr::str_replace('NA', NA_character_) + unique_probe_ids$Gene.Title <- purrr::map_chr(stringr::str_split(unique_probe_ids$Gene.Title, stringr::fixed(' /// ')), ~paste0(unique(.), collapse = "|")) %>% stringr::str_replace('NA', NA_character_) + unique_probe_ids$Entrez.Gene <- purrr::map_chr(stringr::str_split(unique_probe_ids$Entrez.Gene, stringr::fixed(' /// ')), ~paste0(unique(.), collapse = "|")) %>% stringr::str_replace('NA', NA_character_) + unique_probe_ids$Ensembl <- purrr::map_chr(stringr::str_split(unique_probe_ids$Ensembl, stringr::fixed(' /// ')), ~paste0(unique(.), collapse = "|")) %>% stringr::str_replace('NA', NA_character_) + + unique_probe_ids$RefSeq <- paste(unique_probe_ids$RefSeq.Transcript.ID, unique_probe_ids$RefSeq.Protein.ID) + unique_probe_ids$RefSeq <- purrr::map_chr(stringr::str_extract_all(unique_probe_ids$RefSeq, '[A-Z]+_[\\d.]+'), ~paste0(unique(.), collapse = "|")) %>% stringr::str_replace('^$', NA_character_) + + unique_probe_ids$GO <- paste(unique_probe_ids$Gene.Ontology.Biological.Process, unique_probe_ids$Gene.Ontology.Cellular.Component, unique_probe_ids$Gene.Ontology.Molecular.Function) + unique_probe_ids$GO <- purrr::map_chr(stringr::str_extract_all(unique_probe_ids$GO, '\\d{7}'), ~paste0('GO:', unique(.), collapse = "|")) %>% stringr::str_replace('^GO:$', NA_character_) + + unique_probe_ids <- unique_probe_ids[c('Probe.Set.ID', 'Entrez.Gene', 'Gene.Symbol', 'Gene.Title', 'Ensembl', 'RefSeq', 'GO')] + names(unique_probe_ids) <- c('ProbesetID', 'ENTREZID', 'SYMBOL', 'GENENAME', 'ENSEMBL', 'REFSEQ', 'GOSLIM_IDS') + + unique_probe_ids$STRING_id <- NA_character_ + + gene_col <- 'ENSEMBL' + if (sum(!is.na(unique_probe_ids$ENTREZID)) > sum(!is.na(unique_probe_ids$ENSEMBL))) { + gene_col <- 'ENTREZID' + } + if (sum(!is.na(unique_probe_ids$SYMBOL)) > max(sum(!is.na(unique_probe_ids$ENTREZID)), sum(!is.na(unique_probe_ids$ENSEMBL)))) { + gene_col <- 'SYMBOL' + } + + unique_probe_ids <- unique_probe_ids %>% + dplyr::mutate( + count_gene_mappings = 1 + stringr::str_count(get(gene_col), stringr::fixed("|")), + gene_mapping_source = gene_col + ) + } else if (annot_type == 'custom') { + unique_probe_ids <- read.csv( + file.path(local_annotation_dir, annot_config$annot_filename[[1]]), + header = TRUE, na.strings = c('NA', '') + ) + } else { + annot_cols <- c('ProbesetID', 'ENTREZID', 'SYMBOL', 'GENENAME', 'ENSEMBL', 'REFSEQ', 'GOSLIM_IDS', 'STRING_id', 'count_gene_mappings', 'gene_mapping_source') + unique_probe_ids <- setNames(data.frame(matrix(NA_character_, nrow = 1, ncol = length(annot_cols))), annot_cols) + } +} else { + annot <- read.table( + as.character(annotation_file_path), + sep = "\t", + header = TRUE, + quote = "", + comment.char = "" + ) + + unique_probe_ids <- df_mapping %>% + dplyr::mutate(dplyr::across(!!sym(expected_attribute_name), as.character)) %>% # Ensure probeset ids treated as character type + dplyr::group_by(!!sym(expected_attribute_name)) %>% + dplyr::summarise( + ENSEMBL = list_to_unique_piped_string(ensembl_gene_id) + ) %>% + # Count number of ensembl IDS mapped + dplyr::mutate( + count_gene_mappings = 1 + stringr::str_count(ENSEMBL, stringr::fixed("|")), + gene_mapping_source = annot_key + ) %>% + dplyr::left_join(annot, by = c("ENSEMBL" = annot_key)) +} + +probeset_expression_matrix <- oligo::exprs(probeset_level_data) + +probeset_expression_matrix.gene_mapped <- probeset_expression_matrix %>% + as.data.frame() %>% + tibble::rownames_to_column(var = "ProbesetID") %>% # Ensure rownames (probeset IDs) can be used as join key + dplyr::left_join(unique_probe_ids, by = c("ProbesetID" = expected_attribute_name ) ) %>% + dplyr::mutate( count_gene_mappings := ifelse(is.na(count_gene_mappings), 0, count_gene_mappings) ) %>% + dplyr::mutate( gene_mapping_source := unique(unique_probe_ids$gene_mapping_source) ) +``` + +**Custom Functions Used:** + +- [resolve_mart_ftp_base()](#resolve_mart_ftp_base) +- [download_mart_dump()](#download_mart_dump) +- [get_transcript_to_gene_mapping_from_ftp()](#get_transcript_to_gene_mapping_from_ftp) +- [get_probe_to_gene_mapping_from_ftp()](#get_probe_to_gene_mapping_from_ftp) +- [retry_with_delay()](#retry_with_delay) +- [shortened_organism_name()](#shortened_organism_name) +- [get_biomart_attribute()](#get_biomart_attribute) +- [list_to_unique_piped_string()](#list_to_unique_piped_string) + +**Input Data:** + +- `df_rs$organism` (organism specified in the runsheet created in [Step 1](#1-create-sample-runsheet)) +- `df_rs$biomart_attribute` (array design BioMart identifier specified in the runsheet created in [Step 1](#1-create-sample-runsheet)) +- `annotation_file_path` (reference organism annotation file url indicated in the 'genelab_annots_link' column of the [GL-DPPD-7110-A_annotations.csv](https://github.com/nasa/GeneLab_Data_Processing/blob/master/GeneLab_Reference_Annotations/Pipeline_GL-DPPD-7110_Versions/GL-DPPD-7110-A/GL-DPPD-7110-A_annotations.csv) GeneLab Annotations file) +- `ensembl_version` (reference organism Ensembl version indicated in the 'ensemblVersion' column of the [GL-DPPD-7110-A_annotations.csv](https://github.com/nasa/GeneLab_Data_Processing/blob/master/GeneLab_Reference_Annotations/Pipeline_GL-DPPD-7110_Versions/GL-DPPD-7110-A/GL-DPPD-7110-A_annotations.csv) GeneLab Annotations file) +- `annot_key` (keytype to join annotation table and microarray probes, dependent on organism, e.g. mus musculus uses 'ENSEMBL') +- `local_annotation_dir` (Path to local annotation directory if using custom annotations) + > [!TIP] + > If not using custom annotations, leave `local_annotation_dir` as `NULL`. +- `array_annot_path` (URL or path to array design info file if using custom annotations) + > [!TIP] + > If not using custom annotations, leave `array_annot_path` as `NULL`. + + > [!TIP] + > See [design_info.csv](../Workflow_Documentation/NF_MAAffymetrix/examples/annotations/design_info.csv) for the latest array design info file used at GeneLab. This file can also be created manually by following the [file specification](../Workflow_Documentation/NF_MAAffymetrix/examples/annotations/README.md). + +- `probeset_level_data` (R object containing probeset level expression values after summarization of normalized probeset level data, output from [Step 7](#7-probeset-summarization)) + +**Output Data:** + +- `unique_probe_ids` (R object containing probeset ID to gene annotation mappings) +- `probeset_expression_matrix.gene_mapped` (R object containing probeset level expression values after summarization of normalized probeset level data combined with gene annotations specified by Ensembl FTP mart dumps or custom annotations) + +
+ +### 8b. Summarize Gene Mapping + +```R +# Pie Chart with Percentages +slices <- c( + 'Unique Mapping' = nrow(probeset_expression_matrix.gene_mapped %>% dplyr::filter(count_gene_mappings == 1) %>% dplyr::distinct(ProbesetID)), + 'Multi Mapping' = nrow(probeset_expression_matrix.gene_mapped %>% dplyr::filter(count_gene_mappings > 1) %>% dplyr::distinct(ProbesetID)), + 'No Mapping' = nrow(probeset_expression_matrix.gene_mapped %>% dplyr::filter(count_gene_mappings == 0) %>% dplyr::distinct(ProbesetID)) +) +pct <- round(slices/sum(slices)*100) +chart_names <- names(slices) +chart_names <- glue::glue("{names(slices)} ({slices})") # add count to labels +chart_names <- paste(chart_names, pct) # add percents to labels +chart_names <- paste(chart_names,"%",sep="") # ad % to labels +pie(slices,labels = chart_names, col=rainbow(length(slices)), + main=glue::glue("Mapping to Primary Keytype\n {nrow(probeset_expression_matrix.gene_mapped %>% dplyr::distinct(ProbesetID))} Total Unique Probesets") + ) + +print(glue::glue("Unique Mapping Count: {slices[['Unique Mapping']]}")) +``` + +**Input Data:** + +- `probeset_expression_matrix.gene_mapped` (R object containing probeset level expression values after summarization of normalized probeset level data combined with gene annotations specified by Ensembl FTP mart dumps or custom annotations, output from [Step 8a](#8a-get-probeset-annotations) above) + +**Output Data:** + +- A pie chart denoting the gene mapping rates for each unique probeset ID +- A printout denoting the count of unique mappings for gene mapping + +
+ +### 8c. Generate Annotated Raw and Normalized Expression Tables + +```R +## Reorder columns before saving to file +ANNOTATIONS_COLUMN_ORDER = c( + annot_key, + "SYMBOL", + "GENENAME", + "REFSEQ", + "ENTREZID", + "STRING_id", + "GOSLIM_IDS" +) + +SAMPLE_COLUMN_ORDER <- df_rs$`Sample Name` + +probeset_expression_matrix.gene_mapped <- probeset_expression_matrix.gene_mapped %>% dplyr::rename( !!annot_key := ENSEMBL ) + +## Output column subset file with just normalized probeset level expression values +write.csv( + probeset_expression_matrix.gene_mapped[c( + ANNOTATIONS_COLUMN_ORDER, + "ProbesetID", + "count_gene_mappings", + "gene_mapping_source", + SAMPLE_COLUMN_ORDER) + ], file.path(DIR_NORMALIZED_EXPRESSION, "normalized_expression_probeset_GLmicroarray.csv"), row.names = FALSE) + +## Determine column order for probe level tables + +PROBE_INFO_COLUMN_ORDER = c( + "ProbesetID", + "ProbeID", + "count_gene_mappings", + "gene_mapping_source" +) + +FINAL_COLUMN_ORDER <- c( + ANNOTATIONS_COLUMN_ORDER, + PROBE_INFO_COLUMN_ORDER, + SAMPLE_COLUMN_ORDER +) + +## Generate raw intensity matrix that includes annotations + +background_corrected_data_annotated <- oligo::exprs(background_corrected_data) %>% + as.data.frame() %>% + tibble::rownames_to_column(var = "fid") %>% # Ensure rownames (probeset IDs) can be used as join key + dplyr::mutate(dplyr::across(fid, as.integer)) %>% # Ensure fid is integer type, consistent with getProbeInfo typing + dplyr::right_join(oligo::getProbeInfo(background_corrected_data), by = "fid") %>% # Add 'man_fsetid' via mapping based on fid + dplyr::rename( ProbesetID = man_fsetid ) %>% # Rename from getProbeInfo name to ProbesetID + dplyr::rename( ProbeID = fid ) %>% # Rename from getProbeInfo name to ProbeID + dplyr::left_join(unique_probe_ids, by = c("ProbesetID" = expected_attribute_name ) ) %>% # Join with ENSEMBL mappings + dplyr::mutate( count_gene_mappings := ifelse(is.na(count_gene_mappings), 0, count_gene_mappings) ) %>% # Convert NA mapping to 0 + dplyr::mutate( gene_mapping_source := unique(unique_probe_ids$gene_mapping_source) ) %>% + dplyr::rename( !!annot_key := ENSEMBL ) + +## Perform reordering +background_corrected_data_annotated <- background_corrected_data_annotated %>% + dplyr::relocate(dplyr::all_of(FINAL_COLUMN_ORDER)) + +write.csv(background_corrected_data_annotated, file.path(DIR_RAW_DATA, "raw_intensities_probe_GLmicroarray.csv"), row.names = FALSE) + +## Generate normalized expression matrix that includes annotations +norm_data_matrix_annotated <- oligo::exprs(norm_data) %>% + as.data.frame() %>% + tibble::rownames_to_column(var = "fid") %>% # Ensure rownames (probeset IDs) can be used as join key + dplyr::mutate(dplyr::across(fid, as.integer)) %>% # Ensure fid is integer type, consistent with getProbeInfo typing + dplyr::right_join(oligo::getProbeInfo(norm_data), by = "fid") %>% # Add 'man_fsetid' via mapping based on fid + dplyr::rename( ProbesetID = man_fsetid ) %>% # Rename from getProbeInfo name to ProbesetID + dplyr::rename( ProbeID = fid ) %>% # Rename from getProbeInfo name to ProbeID + dplyr::left_join(unique_probe_ids, by = c("ProbesetID" = expected_attribute_name ) ) %>% + dplyr::mutate( count_gene_mappings := ifelse(is.na(count_gene_mappings), 0, count_gene_mappings) ) %>% # Convert NA mapping to 0 + dplyr::mutate( gene_mapping_source := unique(unique_probe_ids$gene_mapping_source) ) %>% + dplyr::rename( !!annot_key := ENSEMBL ) + +norm_data_matrix_annotated <- norm_data_matrix_annotated %>% + dplyr::relocate(dplyr::all_of(FINAL_COLUMN_ORDER)) + +write.csv(norm_data_matrix_annotated, file.path(DIR_NORMALIZED_EXPRESSION, "normalized_intensities_probe_GLmicroarray.csv"), row.names = FALSE) +``` + +**Input Data:** + +- `df_rs` (R dataframe containing information from the runsheet, output from [Step 2c](#2c-load-metadata-and-raw-data)) +- `annot_key` (keytype to join annotation table and microarray probes, dependent on organism, e.g. mus musculus uses 'ENSEMBL', defined in [Step 8a](#8a-get-probeset-annotations)) +- `probeset_expression_matrix.gene_mapped` (R object containing probeset level expression values after summarization of normalized probeset level data combined with gene annotations specified by Ensembl FTP mart dumps or custom annotations, output from [Step 8a](#8a-get-probeset-annotations) above) +- `background_corrected_data` (R object containing background-corrected microarray data, output from [Step 4](#4-background-correction)) +- `norm_data` (R object containing background-corrected and normalized microarray data, output from [Step 5](#5-between-array-normalization)) +- `unique_probe_ids` (R object containing probeset ID to gene annotation mappings, output from [Step 8a](#8a-get-probeset-annotations)) + +**Output Data:** + +- **normalized_expression_probeset_GLmicroarray.csv** (table containing the background corrected, normalized probeset expression values for each sample. The ProbesetID is the unique index column.) +- **raw_intensities_probe_GLmicroarray.csv** (table containing the background corrected, unnormalized probe intensity values for each sample including gene annotations. The ProbeID is the unique index column.) +- **normalized_intensities_probe_GLmicroarray.csv** (table containing the background corrected, normalized probe intensity values for each sample including gene annotations. The ProbeID is the unique index column.) + +## 9. Probeset Differential Expression (DE) + +> [!WARNING] +> Run differential expression analysis only if there are at least 2 replicates per factor group. + +
+ +### 9a. Generate Design Matrix + +```R +# Loading metadata from runsheet csv file +design_data <- runsheet_to_design_matrix(runsheet) +design <- design_data$matrix + +# Write SampleTable.csv and contrasts.csv file +write.csv(design_data$groups, file.path(DIR_DGE, "SampleTable_GLmicroarray.csv"), row.names = FALSE) +write.csv(design_data$contrasts, file.path(DIR_DGE, "contrasts_GLmicroarray.csv")) +``` + +**Custom Functions Used:** + +- [runsheet_to_design_matrix()](#runsheet_to_design_matrix) + +**Input Data:** + +- `runsheet` (path to runsheet, output from [Step 1](#1-create-sample-runsheet)) + +**Output Data:** + +- `design_data` (a list of R objects containing the sample information and metadata + - `design_data$matrix` - the limma study design matrix, indicating the group that each sample belongs to + - `design_data$mapping` - a dataframe of conditions and group names + - `design_data$groups` - a dataframe of samples and group names + - `design_data$contrasts` - a matrix of all pairwise comparisons of the groups) +- `design` (R object containing the limma study design matrix, indicating the group that each sample belongs to) +- **SampleTable_GLmicroarray.csv** (table containing samples and their respective groups) +- **contrasts_GLmicroarray.csv** (table containing all pairwise comparisons) + +
+ +### 9b. Perform Individual Probeset Level DE + +```R +# Calculate results +res <- lm_fit_pairwise(probeset_level_data, design) + +# Print DE table, without filtering +limma::write.fit(res, adjust = 'BH', + file = "INTERIM.csv", + row.names = FALSE, + quote = TRUE, + sep = ",") +``` + +**Custom Functions Used:** + +- [lm_fit_pairwise()](#lm_fit_pairwise) + +**Input Data:** + +- `probeset_level_data` (R object containing probeset level expression values after summarization of normalized probeset level data, output from [Step 7](#7-probeset-summarization)) +- `design` (R object containing the limma study design matrix, indicating the group that each sample belongs to, output from [Step 9a](#9a-generate-design-matrix) above) + +**Output Data:** + +- INTERIM.csv (statistical values from individual probeset level DE analysis, including: + - Log2fc between all pairwise comparisons + - T statistic for all pairwise comparison tests + - P value for all pairwise comparison tests + - Adjusted P value for all pairwise comparison tests) + +
+ +### 9c. Add Annotation and Stats Columns and Format DE Table + +```R +## Reformat Table for consistency across DE analyses tables within GeneLab ## + +# Read in DE table +df_interim <- read.csv("INTERIM.csv") + +# Bind columns from gene mapped expression table +df_interim <- df_interim %>% + dplyr::bind_cols(probeset_expression_matrix.gene_mapped) + +df_interim <- df_interim %>% dplyr::rename_with(reformat_names, .cols = matches('\\.condition'), group_name_mapping = design_data$mapping) + + +## Add Group Wise Statistics ## + +# Group mean and standard deviations for normalized expression values are computed and added to the table + +unique_groups <- unique(design_data$group$group) +for ( i in seq_along(unique_groups) ) { + current_group <- unique_groups[i] + current_samples <- design_data$group %>% + dplyr::filter(group == current_group) %>% + dplyr::pull(sample) %>% + sort() + + print(glue::glue("Computing mean and standard deviation for Group {i} of {length(unique_groups)}")) + print(glue::glue("Group: {current_group}")) + print(glue::glue("Samples in Group: '{toString(current_samples)}'")) + + df_interim <- df_interim %>% + dplyr::mutate( + "Group.Mean_{current_group}" := rowMeans(dplyr::select(., all_of(current_samples))), + "Group.Stdev_{current_group}" := matrixStats::rowSds(as.matrix(dplyr::select(., all_of(current_samples)))), + ) %>% + dplyr::ungroup() %>% + as.data.frame() +} + +df_interim <- df_interim %>% + dplyr::mutate( + "All.mean" := rowMeans(dplyr::select(., all_of(SAMPLE_COLUMN_ORDER))), + "All.stdev" := matrixStats::rowSds(as.matrix(dplyr::select(., all_of(SAMPLE_COLUMN_ORDER)))), + ) %>% + dplyr::ungroup() %>% + as.data.frame() + +print("Remove extra columns from final table") + +# These columns are data mapped to column PROBEID as per the original Manufacturer and can be linked as needed +colnames_to_remove = c( + "AveExpr" # Replaced by 'All.mean' column +) + +df_interim <- df_interim %>% dplyr::select(-any_of(colnames_to_remove)) + +PROBE_INFO_COLUMN_ORDER = c( + "ProbesetID", + "count_gene_mappings", + "gene_mapping_source" +) + +STAT_COLUMNS_ORDER <- generate_prefixed_column_order( + subjects = colnames(design_data$contrasts), + prefixes = c( + "Log2fc_", + "T.stat_", + "P.value_", + "Adj.p.value_" + ) + ) +ALL_SAMPLE_STATS_COLUMNS_ORDER <- c( + "All.mean", + "All.stdev", + "F", + "F.p.value" +) + +GROUP_MEAN_STDEV_COLUMNS_ORDER <- generate_prefixed_column_order( + subjects = unique(design_data$groups$group), + prefixes = c( + "Group.Mean_", + "Group.Stdev_" + ) +) + +FINAL_COLUMN_ORDER <- c( + ANNOTATIONS_COLUMN_ORDER, + PROBE_INFO_COLUMN_ORDER, + SAMPLE_COLUMN_ORDER, + STAT_COLUMNS_ORDER, + ALL_SAMPLE_STATS_COLUMNS_ORDER, + GROUP_MEAN_STDEV_COLUMNS_ORDER +) + +## Assert final column order includes all columns from original table +if (!setequal(FINAL_COLUMN_ORDER, colnames(df_interim))) { + write.csv(FINAL_COLUMN_ORDER, "FINAL_COLUMN_ORDER.csv") + NOT_IN_DF_INTERIM <- paste(setdiff(FINAL_COLUMN_ORDER, colnames(df_interim)), collapse = ":::") + NOT_IN_FINAL_COLUMN_ORDER <- paste(setdiff(colnames(df_interim), FINAL_COLUMN_ORDER), collapse = ":::") + stop(glue::glue("Column reordering attempt resulted in different sets of columns than original. Names unique to 'df_interim': {NOT_IN_FINAL_COLUMN_ORDER}. Names unique to 'FINAL_COLUMN_ORDER': {NOT_IN_DF_INTERIM}.")) +} + +## Perform reordering +df_interim <- df_interim %>% dplyr::relocate(dplyr::all_of(FINAL_COLUMN_ORDER)) + +# Save to file +write.csv(df_interim, file.path(DIR_DGE, "differential_expression_GLmicroarray.csv"), row.names = FALSE) +``` + +**Custom Functions Used:** + +- [reformat_names()](#reformat_names) +- [generate_prefixed_column_order()](#generate_prefixed_column_order) + +**Input Data:** + +- `design_data` (a list of R objects containing the sample information and metadata, output from [Step 9a](#9a-generate-design-matrix) above) +- INTERIM.csv (statistical values from individual probeset level DE analysis, output from [Step 9b](#9b-perform-individual-probeset-level-de) above) +- `probeset_expression_matrix.gene_mapped` (R object containing probeset level expression values after summarization of normalized probeset level data combined with gene annotations specified by Ensembl FTP mart dumps or custom annotations, output from [Step 8a](#8a-get-probeset-annotations) above) + +**Output Data:** + +- **differential_expression_GLmicroarray.csv** (table containing normalized probeset expression values for each sample, group statistics, Limma probeset DE results for each pairwise comparison, and gene annotations. The ProbesetID is the unique index column.) + +
+ +--- + +> [!IMPORTANT] +> All steps of the Microarray pipeline are performed using R markdown and the completed R markdown is rendered (via Quarto) as an html file (**NF_MAAffymetrix_v\*_GLmicroarray.html**) and published in the [Open Science Data Repository (OSDR)](https://osdr.nasa.gov/bio/repo/) for the respective dataset. diff --git a/Microarray/Affymetrix/README.md b/Microarray/Affymetrix/README.md index 6743ea3ae..5994609bb 100644 --- a/Microarray/Affymetrix/README.md +++ b/Microarray/Affymetrix/README.md @@ -1,7 +1,7 @@ # GeneLab bioinformatics processing pipeline for Affymetrix microarray data -> **The document [`GL-DPPD-7114.md`](Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114.md) holds an overview and example commands for how GeneLab processes Affymetrix microarray datasets. See the [Repository Links](#repository-links) descriptions below for more information. Processed data output files and processing code is provided for each GLDS dataset along with the processed data in the [Open Science Data Repository (OSDR)](https://osdr.nasa.gov/bio/repo/).** +> **The document [`GL-DPPD-7114-A.md`](Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114-A.md) holds an overview and example commands for how GeneLab processes Affymetrix microarray datasets. See the [Repository Links](#repository-links) descriptions below for more information. Processed data output files and processing code is provided for each GLDS dataset along with the processed data in the [Open Science Data Repository (OSDR)](https://osdr.nasa.gov/bio/repo/).** --- diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/CHANGELOG.md b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/CHANGELOG.md index de658e22b..78027f04c 100644 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/CHANGELOG.md +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/CHANGELOG.md @@ -5,24 +5,47 @@ All notable changes to this project will be documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). -## [1.0.5](https://github.com/nasa/GeneLab_Data_Processing/tree/NF_MAAffymetrix_1.0.5/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix) - 2024-08-30 +## [1.0.5](https://github.com/nasa/GeneLab_Data_Processing/tree/NF_MAAffymetrix_1.0.5/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix) - 2026-05-XX ### Added -- Add support for bacteria annotations using manufacturer annotations ([#113](https://github.com/nasa/GeneLab_Data_Processing/issues/113)) +- Support for custom annotations, see [specification](examples/annotations/README.md) ([#113](https://github.com/nasa/GeneLab_Data_Processing/issues/113)) - Add option to skip differential expression analysis (`--skipDE`) ([#104](https://github.com/nasa/GeneLab_Data_Processing/issues/104)) +- Add nextflow schema support for parameter validation and help text generation +- Add conda support for easier local development and debugging ### Changed -- Small bug fixes in `Affymetrix.qmd` - - Check if `getBM()` returned results before concatenating it to dataframe to avoid error in `bind_rows()` ([#96](https://github.com/nasa/GeneLab_Data_Processing/issues/96)) +- Replace `RUNSHEET_FROM_GLDS` and `RUNSHEET_FROM_ISA` processes and their associated workflow logic with a new staging analysis subworkflow supporting both accession-based and input-file-based execution modes +- Rework publish directory behavior as part of the staging analysis subworkflow where `outdir` is now the base directory for `GLDS-NNN/` output directory if `--accession` is provided, or the base directory for `results/` output directory if `--runsheet` is provided +- Bump gl-microarray image from version 1.0.0 to 1.1.0 to match R package updates in the [GL-DPPD-7114-A pipeline document](../../Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114-A.md) +- Rename `annotation_config_path` as `array_annot_path` and `config.csv` as `design_info.csv` throughout the workflow and documentation to better reflect the purpose of the file and its contents +- Convert `generate_protocol.sh` to a Python script for automated handling of reference/annotation parameters +- Add `create_date` to design_info.csv and parse it in the protocol +- Move protocol creation from post-processing to main nextflow script to make passing needed values easier and more robust +- Update software table generation to exclude `purrr` from table if custom annotations are not used +- Rename module files from UPPERCASE.nf to lowercase.nf following Nextflow community convention +- Flatten directory-based modules (PROCESS_NAME/ with scripts under resources/usr/bin/) to single lowercase process_name.nf files directly under modules/ +- Move process scripts from modules/PROCESS_NAME/resources/usr/bin/ to the top-level bin/ directory +- Update processed data protocol to auto-populate workflow version from `nextflow.config` and add Caenorhabditis elegans, Saccharomyces cerevisiae, Escherichia coli, and Pseudomonas aeruginosa to supported organisms ([#98](https://github.com/nasa/GeneLab_Data_Processing/issues/98)) +- Fixes in `Affymetrix.qmd` + - Replace live biomaRt::getBM() queries in the QMD with direct downloads of Ensembl's FTP mart-dump tables; drop chunking/retry/Sys.sleep tied to those queries - When renaming column names, specify which columns to rename to avoid unintentional renaming ([#97](https://github.com/nasa/GeneLab_Data_Processing/issues/97)) - When renaming factor names, prevent cases where a factor is partially renamed because it contains a substring that is another factor ([#100](https://github.com/nasa/GeneLab_Data_Processing/issues/100)) - Update MA plot to support HTAFeatureSet ([#105](https://github.com/nasa/GeneLab_Data_Processing/issues/105)) - Remove extra `.1` suffix in AFFY HTA 2 0 Probe IDs in the raw data to allow for merging to BioMart data ([#106](https://github.com/nasa/GeneLab_Data_Processing/issues/106)) - Decrease legend size when sample names are long to prevent it from covering plot ([#107](https://github.com/nasa/GeneLab_Data_Processing/issues/107)) -- Update processed data protocol to auto-populate workflow version from `nextflow.config` and add Caenorhabditis elegans, Saccharomyces cerevisiae, Escherichia coli, and Pseudomonas aeruginosa to supported organisms ([#98](https://github.com/nasa/GeneLab_Data_Processing/issues/98)) -- Update software table generation to exclude `R.utils` from table if data files are not compressed ([#99](https://github.com/nasa/GeneLab_Data_Processing/issues/99)) + - Simplify group sample retrieval during differential expression group-wise statistics computation to use a more concise `filter/pull/sort` chain instead of `group_by/summarize/filter/pull`, addressing the deprecation warning in dplyr >= 1.1.0 where returning more than 1 row per `summarise()` group is deprecated +- Changes to post-processing workflow + - Resolve output directory `GLDS-NNN/` or `results/` to match main workflow behavior + - Replace dp_tools dependency in assay table update and md5sum table generation with standalone scripts + - Rename `UPDATE_ISA_TABLES` and `update_curation_table.py` to `UPDATE_ASSAY_TABLE` and `update_assay_table.py` to better reflect their purpose + - Add new PURGE_PROCESSING_INFO Nextflow module to strip full paths in nextflow_processing_info_GLmicroarray.txt before publishing + - Add parameter validation and summary log from nf-schema + +### Removed + +- Packages `R.utils`, and `biomaRt` are no longer used in the processing code, and have been removed from software table generation ## [1.0.4](https://github.com/nasa/GeneLab_Data_Processing/tree/NF_MAAffymetrix_1.0.4/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix) - 2024-05-17 diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/README.md b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/README.md index f21d727a4..206f7e381 100644 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/README.md +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/README.md @@ -4,7 +4,7 @@ ### Implementation Tools -The current GeneLab Affymetrix Microarray consensus processing pipeline (NF_MAAffymetrix), [GL-DPPD-7114](../../Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114.md), is implemented as a [Nextflow](https://nextflow.io/) DSL2 workflow and utilizes [Singularity](https://docs.sylabs.io/guides/3.10/user-guide/introduction.html) to run all tools in containers. This workflow (NF_MAAffymetrix) is run using the command line interface (CLI) of any unix-based system. While knowledge of creating workflows in Nextflow is not required to run the workflow as is, [the Nextflow documentation](https://nextflow.io/docs/latest/index.html) is a useful resource for users who want to modify and/or extend this workflow. +The current GeneLab Affymetrix Microarray consensus processing pipeline (NF_MAAffymetrix), [GL-DPPD-7114-A](../../Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114-A.md), is implemented as a [Nextflow](https://nextflow.io/) DSL2 workflow and utilizes [Singularity](https://docs.sylabs.io/guides/3.10/user-guide/introduction.html) or [Docker](https://docs.docker.com/get-started/) containers to run all tools. This workflow (NF_MAAffymetrix) is run using the command line interface (CLI) of any unix-based system. While knowledge of creating workflows in Nextflow is not required to run the workflow as is, [the Nextflow documentation](https://docs.seqera.io/nextflow/) is a useful resource for users who want to modify and/or extend this workflow. ### Workflow & Subworkflows @@ -14,8 +14,8 @@ The current GeneLab Affymetrix Microarray consensus processing pipeline (NF_MAAf --- The NF_MAAffymetrix workflow is composed of three subworkflows as shown in the image above. -Below is a description of each subworkflow and the additional output files generated that are not already indicated in the [GL-DPPD-7114 pipeline -document](../../Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114.md): +Below is a description of each subworkflow and the additional output files generated that are not already indicated in the [GL-DPPD-7114-A pipeline +document](../../Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114-A.md): 1. **Analysis Staging Subworkflow** @@ -26,7 +26,7 @@ document](../../Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114.md): 2. **Affymetrix Microarray Processing Subworkflow** - Description: - - This subworkflow uses the staged raw data and metadata parameters from the Analysis Staging Subworkflow to generate processed data using the [GL-DPPD-7114 pipeline](../../Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114.md). + - This subworkflow uses the staged raw data and metadata parameters from the Analysis Staging Subworkflow to generate processed data using the [GL-DPPD-7114-A pipeline](../../Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114-A.md). 1. **V&V Pipeline Subworkflow** @@ -34,13 +34,13 @@ document](../../Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114.md): - This subworkflow performs validation and verification (V&V) on the raw and processed data files. It performs a series of checks on the output files generated and flags the results, using the flag codes indicated in the table below, which are outputted into a log file. **V&V Flags**: - |Flag Codes|Flag Name|Interpretation| - |:---------|:--------|:-------------| - | 2 | MANUAL | Special flag that indicates a manual check that is advised. Often used to advise what should be visually assessed in QA plots. | - | 20 | GREEN | Indicates the check passed all validation conditions | - | 30 | YELLOW | Indicates the check was flagged for minor issues (e.g. slight outliers) | - | 50 | RED | Indicates the check was flagged for moderate issues (e.g. major outliers) | - | 80 | HALT | Indicates the check was flagged for severe issues that trigger a processing halt (e.g. missing data) | + | Flag Codes | Flag Name | Interpretation | + | :--------- | :-------- | :----------------------------------------------------------------------------------------------------------------------------- | + | 2 | MANUAL | Special flag that indicates a manual check that is advised. Often used to advise what should be visually assessed in QA plots. | + | 20 | GREEN | Indicates the check passed all validation conditions | + | 30 | YELLOW | Indicates the check was flagged for minor issues (e.g. slight outliers) | + | 50 | RED | Indicates the check was flagged for moderate issues (e.g. major outliers) | + | 80 | HALT | Indicates the check was flagged for severe issues that trigger a processing halt (e.g. missing data) |
@@ -66,9 +66,10 @@ document](../../Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114.md): #### 1a. Install Nextflow -Nextflow can be installed either through [Anaconda](https://anaconda.org/bioconda/nextflow) or as documented on the [Nextflow documentation page](https://www.nextflow.io/docs/latest/getstarted.html). +Nextflow can be installed either through the [Anaconda bioconda channel](https://anaconda.org/bioconda/nextflow) or as documented in the [Nextflow installation documentation](https://docs.seqera.io/nextflow/install). -> Note: If you want to install Anaconda, we recommend installing a Miniconda, Python3 version appropriate for your system, as instructed by [Happy Belly Bioinformatics](https://astrobiomike.github.io/unix/conda-intro#getting-and-installing-conda). +> [!TIP] +> If you wish to install Anaconda, we recommend installing a Miniforge version appropriate for your system, as documented on the [conda-forge website](https://conda-forge.org/download/), where you can find basic binaries for most systems. More detailed miniforge documentation is available in the [miniforge github repository](https://github.com/conda-forge/miniforge). > > Once conda is installed on your system, you can install the latest version of Nextflow by running the following commands: > @@ -85,7 +86,9 @@ Singularity is a container platform that allows usage of containerized software. We recommend installing Singularity on a system wide level as per the associated [documentation](https://docs.sylabs.io/guides/3.10/admin-guide/admin_quickstart.html). -> Note: Singularity is also available through [Anaconda](https://anaconda.org/conda-forge/singularity). +> [!TIP] +> - Singularity is also available through the [Anaconda conda-forge channel](https://anaconda.org/conda-forge/singularity). +> - Alternatively, Docker can be used in place of Singularity. To get started with Docker, see the [Docker CE installation documentation](https://docs.docker.com/engine/install/).
@@ -109,7 +112,8 @@ unzip NF_MAAffymetrix_1.0.5.zip ### 3. Run the Workflow While in the location containing the `NF_MAAffymetrix_1.0.5` directory that was downloaded in [step 2](#2-download-the-workflow-files), you are now able to run the workflow. Below are three examples of how to run the NF_MAAffymetrix workflow: -> Note: Nextflow commands use both single hyphen arguments (e.g. -help) that denote general nextflow arguments and double hyphen arguments (e.g. --ensemblVersion) that denote workflow specific parameters. Take care to use the proper number of hyphens for each argument. +> [!NOTE] +> Nextflow commands use both single hyphen arguments (e.g. -help) that denote general nextflow arguments and double hyphen arguments (e.g. --ensemblVersion) that denote workflow specific parameters. Take care to use the proper number of hyphens for each argument.
@@ -118,15 +122,15 @@ While in the location containing the `NF_MAAffymetrix_1.0.5` directory that was ```bash nextflow run NF_MAAffymetrix_1.0.5/main.nf \ -profile singularity \ - --osdAccession OSD-266 \ - --gldsAccession GLDS-266 + --accession OSD-266 ```
#### 3b. Approach 2: Run the workflow on a non-GLDS dataset using a user-created runsheet -> Note: Specifications for creating a runsheet manually are described [here](examples/runsheet/README.md). +> [!NOTE] +> Specifications for creating a runsheet manually are described [here](examples/runsheet/README.md). ```bash nextflow run NF_MAAffymetrix_1.0.5/main.nf \ @@ -138,11 +142,13 @@ nextflow run NF_MAAffymetrix_1.0.5/main.nf \ #### 3c. Approach 3: Run the workflow using an ISA Archive -> Note: Specifications for the ISA Tab Archive format can be found [here](https://isa-specs.readthedocs.io/en/latest/isatab.html). +> [!NOTE] +> Specifications for the ISA Tab Archive format can be found [here](https://isa-specs.readthedocs.io/en/latest/isatab.html). ```bash nextflow run NF_MAAffymetrix_1.0.5/main.nf \ -profile singularity \ + --accession OSD-266 \ --isaArchivePath
``` @@ -157,11 +163,9 @@ nextflow run NF_MAAffymetrix_1.0.5/main.nf \
-**Additional Required Parameters For [Approach 1](#3a-approach-1-run-the-workflow-on-a-genelab-agilent-1-channel-microarray-dataset):** +**Additional Required Parameters For [Approach 1](#3a-approach-1-run-the-workflow-on-a-genelab-affymetrix-microarray-dataset):** -* `--osdAccession OSD-###` – specifies the OSD ID to process through the NF_MAAffymetrix workflow (replace ### with the OSD number) - -* `--gldsAccession GLDS-###` – specifies the GLDS ID to process through the NF_MAAffymetrix workflow (replace ### with the GLDS number) +* `--accession` – The OSD or GLDS ID for the dataset to be processed, eg. `OSD-266` or `GLDS-266`
@@ -171,11 +175,21 @@ nextflow run NF_MAAffymetrix_1.0.5/main.nf \
+**Additional Required Parameters For [Approach 3](#3c-approach-3-run-the-workflow-using-an-isa-archive):** + +* `--accession` – The OSD or GLDS ID for the dataset to be processed, eg. `OSD-266` or `GLDS-266` + +* `--isaArchivePath` - specifies the path to a previously-downloaded *ISA.zip (Default: an *ISA.zip is automatically fetched from the GeneLab Repository for the GLDS dataset being processed) + +
+ **Optional Parameters:** * `--skipVV` - skip the automated V&V processes (Default: the automated V&V processes are active) -* `--resultsDir` - specifies the output directory for all files produced by the workflow (Default: if OSD and GLDS accessions are specified. Otherwise, the workflow launch directory.) +* `--skipDE` - skip the differential expression analysis (Default: the differential expression analysis is performed) + +* `--outdir` - specifies the base directory where the output directory will be created (Default: "${launchDir}")
@@ -185,7 +199,7 @@ All parameters listed above and additional optional arguments for the NF_MAAffym nextflow run NF_MAAffymetrix_1.0.5/main.nf --help ``` -See `nextflow run -h` and [Nextflow's CLI run command documentation](https://nextflow.io/docs/latest/cli.html#run) for more options and details common to all nextflow workflows. +See `nextflow run -h` and [Nextflow's CLI run command documentation](https://docs.seqera.io/nextflow/cli) for more options and details common to all nextflow workflows.
@@ -193,17 +207,22 @@ See `nextflow run -h` and [Nextflow's CLI run command documentation](https://nex ### 4. Additional Output Files -All R code steps and output are rendered within a Quarto document yielding the following: +> [!NOTE] +> The outputs from the Affymetrix Microarray Processing Subworkflow are documented in the [GL-DPPD-7114-A.md](../../Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114-A.md) processing protocol. + +Additional outputs are described below: - Output: - - NF_MAAffymetrix_1.0.5.html (html report containing executed code and output including QA plots) + - NF_MAAffymetrix_1.0.5_GLmicroarray.html (html report containing executed R code and output including QA plots rendered by Quarto) + - protocol_GLmicroarray.txt (text file describing the processing methods used by the workflow) + - software_versions_GLmicroarray.md (version capturing file for all tools and packages used in the workflow) The outputs from the Analysis Staging and V&V Pipeline Subworkflows are described below: -> Note: The outputs from the Affymetrix Microarray Processing Subworkflow are documented in the [GL-DPPD-7114.md](../../../Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114.md) processing protocol. **Analysis Staging Subworkflow** - +> [!NOTE] +> only applicable for [Approach 1](#3a-approach-1-run-the-workflow-on-a-genelab-affymetrix-microarray-dataset) and [Approach 3](#3c-approach-3-run-the-workflow-using-an-isa-archive) - Output: - \*_microarray_v1_runsheet.csv (table containing metadata required for processing, including the raw reads files location) - \*-ISA.zip (the ISA archive of the GLDS datasets to be processed, downloaded from the GeneLab Data Repository) @@ -212,12 +231,12 @@ The outputs from the Analysis Staging and V&V Pipeline Subworkflows are describe **V&V Pipeline Subworkflow** - Output: - - VV_log_VV_AGILE1CH.tsv.MANUAL_CHECKS_PENDING (table containing V&V flags for all checks performed. Also contains rows indicating suggested manual checks focusing on QA plots embedded in the html report) + - VV_log_VV_AFFYMETRIX_GLmicroarray.tsv.MANUAL_CHECKS_PENDING (table containing V&V flags for all checks performed. Also contains rows indicating suggested manual checks focusing on QA plots embedded in the html report)
Standard Nextflow resource usage logs are also produced as follows: -> Further details about these logs can also found within [this Nextflow documentation page](https://www.nextflow.io/docs/latest/tracing.html#execution-report). +> Further details about these logs can also be found in the [Nextflow Report Documentation](https://docs.seqera.io/nextflow/reports). **Nextflow Resource Usage Logs** - Output: diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/examples/annotations/README.md b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/examples/annotations/README.md new file mode 100644 index 000000000..c17a30b3a --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/examples/annotations/README.md @@ -0,0 +1,30 @@ +# Custom Annotations Specification + +## Description + +* If using custom gene annotations when processing Affymetrix datasets through GeneLab's Affymetrix processing pipeline, a CSV design information file must be provided as specified below. +* See [design_info.csv](design_info.csv) for the latest design information file used at GeneLab. + + +## Example + +- [design_info.csv](design_info.csv) + + +## Required columns + +| Column Name | Type | Description | Example | +|:------------|:-----|:------------|:--------| +| array_design | string | A bioMart attribute identifier denoting the microarray probe/probeset attribute used for annotation mapping. | AFFY E coli Genome 2 0 | +| annot_type | string | Used to determine how the custom annotations are parsed before merging to the data. Currently, only the below are supported: | 3prime-IVT | +| annot_filename | string | Name of the custom annotations file. | E_coli_2.na36.annot.csv | + +## Optional columns +If the file was downloaded from a website, provide download link, download date, +and create date (if applicable) in additional columns after the required column for traceability. + +| Column Name | Type | Description | Example | +|:------------|:-----|:------------|:--------| +| download_link | string | The URL used to retrieve the annotation file. | https://www.thermofisher.com/order/catalog/product/sec/assets?url=TFS-Assets/LSG/Support-Files/E_coli_2-na36-annot-csv.zip | +| download_date | date string | The date the file was retrieved in YYYY-MM-DD format. | 2024-06-15 | +| create_date | date string | The date the file was created in YYYY-MM-DD format. | 2024-03-30 | diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/examples/annotations/design_info.csv b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/examples/annotations/design_info.csv new file mode 100644 index 000000000..837badef8 --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/examples/annotations/design_info.csv @@ -0,0 +1,3 @@ +array_design,annot_type,annot_filename,download_link,download_date,create_date +AFFY E coli Genome 2 0,3prime-IVT,E_coli_2.na36.annot.csv,https://www.thermofisher.com/order/catalog/product/sec/assets?url=TFS-Assets/LSG/Support-Files/E_coli_2-na36-annot-csv.zip,2024-06-15,2016-03-30 +AFFY GeneChip P. aeruginosa Genome,3prime-IVT,Pae_G1a.na36.annot.csv,https://www.thermofisher.com/order/catalog/product/sec/assets?url=TFS-Assets/LSG/Support-Files/Pae_G1a-na36-annot-csv.zip,2024-06-15,2016-03-30 diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/examples/runsheet/OSD-213_microarray_v0_runsheet.csv b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/examples/runsheet/OSD-213_microarray_v0_runsheet.csv new file mode 100644 index 000000000..f309c8545 --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/examples/runsheet/OSD-213_microarray_v0_runsheet.csv @@ -0,0 +1,7 @@ +Sample Name,Study Assay Measurement,Study Assay Technology Type,Study Assay Technology Platform,organism,biomart_attribute,Source Name,Label,Array Data File Name,Array Data File Path,Comment[Array Data File Name],Factor Value[Spaceflight],Factor Value[Altered Gravity],Original Sample Name +Atha_Col-0_clsCC_FLT_1G_Rep1,transcription profiling,DNA microarray,Affymetrix,Arabidopsis thaliana,AFFY ATH1 121501,Culture cells_1,biotin,GLDS-213_microarray_FC_front.CEL.gz,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-213/download?source=datamanager&file=GLDS-213_microarray_FC_front.CEL.gz,GLDS-213_microarray_FC_front.CEL.gz,Space Flight,1G by centrifugation,Atha_Col-0_clsCC_FLT_1G_Rep1 +Atha_Col-0_clsCC_FLT_1G_Rep2,transcription profiling,DNA microarray,Affymetrix,Arabidopsis thaliana,AFFY ATH1 121501,Culture cells_2,biotin,GLDS-213_microarray_FC_rear.CEL.gz,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-213/download?source=datamanager&file=GLDS-213_microarray_FC_rear.CEL.gz,GLDS-213_microarray_FC_rear.CEL.gz,Space Flight,1G by centrifugation,Atha_Col-0_clsCC_FLT_1G_Rep2 +Atha_Col-0_clsCC_FLT_uG_Rep1,transcription profiling,DNA microarray,Affymetrix,Arabidopsis thaliana,AFFY ATH1 121501,Culture cells_3,biotin,GLDS-213_microarray_FS_front.CEL.gz,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-213/download?source=datamanager&file=GLDS-213_microarray_FS_front.CEL.gz,GLDS-213_microarray_FS_front.CEL.gz,Space Flight,uG,Atha_Col-0_clsCC_FLT_uG_Rep1 +Atha_Col-0_clsCC_FLT_uG_Rep2,transcription profiling,DNA microarray,Affymetrix,Arabidopsis thaliana,AFFY ATH1 121501,Culture cells_4,biotin,GLDS-213_microarray_FS_rear.CEL.gz,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-213/download?source=datamanager&file=GLDS-213_microarray_FS_rear.CEL.gz,GLDS-213_microarray_FS_rear.CEL.gz,Space Flight,uG,Atha_Col-0_clsCC_FLT_uG_Rep2 +Atha_Col-0_clsCC_GC_1G_Rep1,transcription profiling,DNA microarray,Affymetrix,Arabidopsis thaliana,AFFY ATH1 121501,Culture cells_5,biotin,GLDS-213_microarray_GS_front.CEL.gz,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-213/download?source=datamanager&file=GLDS-213_microarray_GS_front.CEL.gz,GLDS-213_microarray_GS_front.CEL.gz,Ground Control,1G on Earth,Atha_Col-0_clsCC_GC_1G_Rep1 +Atha_Col-0_clsCC_GC_1G_Rep2,transcription profiling,DNA microarray,Affymetrix,Arabidopsis thaliana,AFFY ATH1 121501,Culture cells_6,biotin,GLDS-213_microarray_GS_rear.CEL.gz,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-213/download?source=datamanager&file=GLDS-213_microarray_GS_rear.CEL.gz,GLDS-213_microarray_GS_rear.CEL.gz,Ground Control,1G on Earth,Atha_Col-0_clsCC_GC_1G_Rep2 diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/examples/runsheet/OSD-3_microarray_v0_runsheet.csv b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/examples/runsheet/OSD-3_microarray_v0_runsheet.csv new file mode 100644 index 000000000..fdc974dee --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/examples/runsheet/OSD-3_microarray_v0_runsheet.csv @@ -0,0 +1,19 @@ +Sample Name,Study Assay Measurement,Study Assay Technology Type,Study Assay Technology Platform,organism,biomart_attribute,Source Name,Label,Array Data File Name,Array Data File Path,Comment[Array Data File Name],Factor Value[Developmental Stage],Factor Value[Spaceflight],Original Sample Name +Dmel_Hml-GAL4-UAS-GFP_wo_FLT_3rd-Instar-Larva_Rep1,transcription profiling,DNA microarray,Affymetrix,Drosophila melanogaster,AFFY Drosophila 2,GSM588931 1,biotin,GLDS-3_microarray_GSM588948.CEL,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-3/download?source=datamanager&file=GLDS-3_microarray_GSM588948.CEL,GLDS-3_microarray_GSM588948.CEL,third instar larva stage,Space Flight,Dmel_Hml-GAL4-UAS-GFP_wo_FLT_3rd-Instar-Larva_Rep1 +Dmel_Hml-GAL4-UAS-GFP_wo_FLT_3rd-Instar-Larva_Rep2,transcription profiling,DNA microarray,Affymetrix,Drosophila melanogaster,AFFY Drosophila 2,GSM588932 1,biotin,GLDS-3_microarray_GSM588947.CEL,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-3/download?source=datamanager&file=GLDS-3_microarray_GSM588947.CEL,GLDS-3_microarray_GSM588947.CEL,third instar larva stage,Space Flight,Dmel_Hml-GAL4-UAS-GFP_wo_FLT_3rd-Instar-Larva_Rep2 +Dmel_Hml-GAL4-UAS-GFP_wo_FLT_3rd-Instar-Larva_Rep3,transcription profiling,DNA microarray,Affymetrix,Drosophila melanogaster,AFFY Drosophila 2,GSM588933 1,biotin,GLDS-3_microarray_GSM588946.CEL,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-3/download?source=datamanager&file=GLDS-3_microarray_GSM588946.CEL,GLDS-3_microarray_GSM588946.CEL,third instar larva stage,Space Flight,Dmel_Hml-GAL4-UAS-GFP_wo_FLT_3rd-Instar-Larva_Rep3 +Dmel_Hml-GAL4-UAS-GFP_wo_FLT_3rd-Instar-Larva_Rep4,transcription profiling,DNA microarray,Affymetrix,Drosophila melanogaster,AFFY Drosophila 2,GSM588934 1,biotin,GLDS-3_microarray_GSM588945.CEL,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-3/download?source=datamanager&file=GLDS-3_microarray_GSM588945.CEL,GLDS-3_microarray_GSM588945.CEL,third instar larva stage,Space Flight,Dmel_Hml-GAL4-UAS-GFP_wo_FLT_3rd-Instar-Larva_Rep4 +Dmel_Hml-GAL4-UAS-GFP_wo_FLT_3rd-Instar-Larva_Rep5,transcription profiling,DNA microarray,Affymetrix,Drosophila melanogaster,AFFY Drosophila 2,GSM588935 1,biotin,GLDS-3_microarray_GSM588944.CEL,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-3/download?source=datamanager&file=GLDS-3_microarray_GSM588944.CEL,GLDS-3_microarray_GSM588944.CEL,third instar larva stage,Space Flight,Dmel_Hml-GAL4-UAS-GFP_wo_FLT_3rd-Instar-Larva_Rep5 +Dmel_Hml-GAL4-UAS-GFP_wo_FLT_3rd-Instar-Larva_Rep6,transcription profiling,DNA microarray,Affymetrix,Drosophila melanogaster,AFFY Drosophila 2,GSM588936 1,biotin,GLDS-3_microarray_GSM588943.CEL,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-3/download?source=datamanager&file=GLDS-3_microarray_GSM588943.CEL,GLDS-3_microarray_GSM588943.CEL,third instar larva stage,Space Flight,Dmel_Hml-GAL4-UAS-GFP_wo_FLT_3rd-Instar-Larva_Rep6 +Dmel_Hml-GAL4-UAS-GFP_wo_FLT_Adult_Rep1,transcription profiling,DNA microarray,Affymetrix,Drosophila melanogaster,AFFY Drosophila 2,GSM588937 1,biotin,GLDS-3_microarray_GSM588942.CEL,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-3/download?source=datamanager&file=GLDS-3_microarray_GSM588942.CEL,GLDS-3_microarray_GSM588942.CEL,adult,Space Flight,Dmel_Hml-GAL4-UAS-GFP_wo_FLT_Adult_Rep1 +Dmel_Hml-GAL4-UAS-GFP_wo_FLT_Adult_Rep2,transcription profiling,DNA microarray,Affymetrix,Drosophila melanogaster,AFFY Drosophila 2,GSM588938 1,biotin,GLDS-3_microarray_GSM588941.CEL,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-3/download?source=datamanager&file=GLDS-3_microarray_GSM588941.CEL,GLDS-3_microarray_GSM588941.CEL,adult,Space Flight,Dmel_Hml-GAL4-UAS-GFP_wo_FLT_Adult_Rep2 +Dmel_Hml-GAL4-UAS-GFP_wo_FLT_Adult_Rep3,transcription profiling,DNA microarray,Affymetrix,Drosophila melanogaster,AFFY Drosophila 2,GSM588939 1,biotin,GLDS-3_microarray_GSM588940.CEL,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-3/download?source=datamanager&file=GLDS-3_microarray_GSM588940.CEL,GLDS-3_microarray_GSM588940.CEL,adult,Space Flight,Dmel_Hml-GAL4-UAS-GFP_wo_FLT_Adult_Rep3 +Dmel_Hml-GAL4-UAS-GFP_wo_GC_3rd-Instar-Larva_Rep1,transcription profiling,DNA microarray,Affymetrix,Drosophila melanogaster,AFFY Drosophila 2,GSM588940 1,biotin,GLDS-3_microarray_GSM588939.CEL,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-3/download?source=datamanager&file=GLDS-3_microarray_GSM588939.CEL,GLDS-3_microarray_GSM588939.CEL,third instar larva stage,Ground Control,Dmel_Hml-GAL4-UAS-GFP_wo_GC_3rd-Instar-Larva_Rep1 +Dmel_Hml-GAL4-UAS-GFP_wo_GC_3rd-Instar-Larva_Rep2,transcription profiling,DNA microarray,Affymetrix,Drosophila melanogaster,AFFY Drosophila 2,GSM588941 1,biotin,GLDS-3_microarray_GSM588938.CEL,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-3/download?source=datamanager&file=GLDS-3_microarray_GSM588938.CEL,GLDS-3_microarray_GSM588938.CEL,third instar larva stage,Ground Control,Dmel_Hml-GAL4-UAS-GFP_wo_GC_3rd-Instar-Larva_Rep2 +Dmel_Hml-GAL4-UAS-GFP_wo_GC_3rd-Instar-Larva_Rep3,transcription profiling,DNA microarray,Affymetrix,Drosophila melanogaster,AFFY Drosophila 2,GSM588942 1,biotin,GLDS-3_microarray_GSM588937.CEL,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-3/download?source=datamanager&file=GLDS-3_microarray_GSM588937.CEL,GLDS-3_microarray_GSM588937.CEL,third instar larva stage,Ground Control,Dmel_Hml-GAL4-UAS-GFP_wo_GC_3rd-Instar-Larva_Rep3 +Dmel_Hml-GAL4-UAS-GFP_wo_GC_3rd-Instar-Larva_Rep4,transcription profiling,DNA microarray,Affymetrix,Drosophila melanogaster,AFFY Drosophila 2,GSM588943 1,biotin,GLDS-3_microarray_GSM588936.CEL,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-3/download?source=datamanager&file=GLDS-3_microarray_GSM588936.CEL,GLDS-3_microarray_GSM588936.CEL,third instar larva stage,Ground Control,Dmel_Hml-GAL4-UAS-GFP_wo_GC_3rd-Instar-Larva_Rep4 +Dmel_Hml-GAL4-UAS-GFP_wo_GC_3rd-Instar-Larva_Rep5,transcription profiling,DNA microarray,Affymetrix,Drosophila melanogaster,AFFY Drosophila 2,GSM588944 1,biotin,GLDS-3_microarray_GSM588935.CEL,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-3/download?source=datamanager&file=GLDS-3_microarray_GSM588935.CEL,GLDS-3_microarray_GSM588935.CEL,third instar larva stage,Ground Control,Dmel_Hml-GAL4-UAS-GFP_wo_GC_3rd-Instar-Larva_Rep5 +Dmel_Hml-GAL4-UAS-GFP_wo_GC_3rd-Instar-Larva_Rep6,transcription profiling,DNA microarray,Affymetrix,Drosophila melanogaster,AFFY Drosophila 2,GSM588945 1,biotin,GLDS-3_microarray_GSM588934.CEL,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-3/download?source=datamanager&file=GLDS-3_microarray_GSM588934.CEL,GLDS-3_microarray_GSM588934.CEL,third instar larva stage,Ground Control,Dmel_Hml-GAL4-UAS-GFP_wo_GC_3rd-Instar-Larva_Rep6 +Dmel_Hml-GAL4-UAS-GFP_wo_GC_Adult_Rep1,transcription profiling,DNA microarray,Affymetrix,Drosophila melanogaster,AFFY Drosophila 2,GSM588946 1,biotin,GLDS-3_microarray_GSM588933.CEL,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-3/download?source=datamanager&file=GLDS-3_microarray_GSM588933.CEL,GLDS-3_microarray_GSM588933.CEL,adult,Ground Control,Dmel_Hml-GAL4-UAS-GFP_wo_GC_Adult_Rep1 +Dmel_Hml-GAL4-UAS-GFP_wo_GC_Adult_Rep2,transcription profiling,DNA microarray,Affymetrix,Drosophila melanogaster,AFFY Drosophila 2,GSM588947 1,biotin,GLDS-3_microarray_GSM588932.CEL,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-3/download?source=datamanager&file=GLDS-3_microarray_GSM588932.CEL,GLDS-3_microarray_GSM588932.CEL,adult,Ground Control,Dmel_Hml-GAL4-UAS-GFP_wo_GC_Adult_Rep2 +Dmel_Hml-GAL4-UAS-GFP_wo_GC_Adult_Rep3,transcription profiling,DNA microarray,Affymetrix,Drosophila melanogaster,AFFY Drosophila 2,GSM588948 1,biotin,GLDS-3_microarray_GSM588931.CEL,https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/OSD-3/download?source=datamanager&file=GLDS-3_microarray_GSM588931.CEL,GLDS-3_microarray_GSM588931.CEL,adult,Ground Control,Dmel_Hml-GAL4-UAS-GFP_wo_GC_Adult_Rep3 diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/examples/runsheet/README.md b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/examples/runsheet/README.md new file mode 100644 index 000000000..c4927a42d --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/examples/runsheet/README.md @@ -0,0 +1,23 @@ +# Runsheet Specification + +## Description + +* The Runsheet is a csv file that contains the metadata required for processing Affymetrix datasets through GeneLab's Affymetrix processing pipeline. + + +## Examples + +1. [Runsheet for GLDS-3](OSD-3_microarray_v0_runsheet.csv) +2. [Runsheet for GLDS-213](OSD-213_microarray_v0_runsheet.csv) + + +## Required columns + +| Column Name | Type | Description | Example | +|:------------|:-----|:------------|:--------| +| Sample Name | string | Sample Name, added as a prefix to sample-specific processed data output files. Should not include spaces or weird characters. | Dmel_Hml-GAL4-UAS-GFP_wo_FLT_3rd-Instar-Larva_Rep1 | +| biomart_attribute | string | A bioMart attribute identifier denoting the microarray probe/probeset attribute used for annotation mapping. | AFFY Drosophila 2 | +| organism | string | Species name used to map to the appropriate gene annotations file. | Drosophila melanogaster | +| Array Data File Path | string (url or local path) | Location of the raw data file for the sample. | /my/data/sample_1.CEL | +| Factor Value[] | string | A set of one or more columns specifying the experimental group the sample belongs to. In the simplest form, a column named 'Factor Value[group]' is sufficient. | Space Flight | +| Original Sample Name | string | Used to map the sample name that will be used for processing to the original sample name. This is often identical except in cases where the original name includes spaces or weird characters. | Dmel_Hml-GAL4-UAS-GFP_wo_FLT_3rd-Instar-Larva_Rep1 | diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/Affymetrix.qmd b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/Affymetrix.qmd index c5f3bec07..53bd0589f 100644 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/Affymetrix.qmd +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/Affymetrix.qmd @@ -1,6 +1,6 @@ --- title: "Affymetrix Processing" -subtitle: "Workflow Version: NF_MAAffymetrix_1.0.5" +subtitle: "`r paste0('Workflow Version: NF_MAAffymetrix_', params$workflow_version)`" date: now title-block-banner: true format: @@ -14,13 +14,14 @@ format: number-sections: true params: + workflow_version: NULL id: NULL # str, used to name output files runsheet: NULL # str, path to runsheet biomart_attribute: NULL # str, used as a fallback value if 'Array Design REF' column is not found in the runsheet - annotation_file_path: NULL # str, Annotation file from 'genelab_annots_link' column of https://github.com/nasa/GeneLab_Data_Processing/blob/GL_RefAnnotTable_1.0.0/GeneLab_Reference_Annotations/Pipeline_GL-DPPD-7110_Versions/GL-DPPD-7110/GL-DPPD-7110_annotations.csv + annotation_file_path: NULL # str, Annotation file from 'genelab_annots_link' column of GeneLab Annotations file ensembl_version: NULL # str, Used to determine ensembl version local_annotation_dir: NULL - DEBUG_limit_biomart_query: NULL # int, If supplied, only the first n probeIDs are queried + array_annot_path: NULL run_DE: 'true' # execute: # DEBUG @@ -44,7 +45,11 @@ if (is.null(params$runsheet)) { stop("PARAMETERIZATION ERROR: Must supply runsheet path") } -runsheet = params$runsheet # +runsheet <- params$runsheet # + +# If using custom annotation, local_annotation_dir is path to directory containing annotation file and array_annot_path is path/url to custom probe annotation file +local_annotation_dir <- params$local_annotation_dir # +array_annot_path <- params$array_annot_path # message(params) @@ -65,7 +70,7 @@ dir.create(DIR_DGE) original_par <- par() options(preferRaster=TRUE) # use Raster when possible to avoid antialiasing artifacts in images -options(timeout=1000) +options(timeout=1000) # ensure enough time for data downloads ``` ## Load Metadata and Raw Data @@ -75,6 +80,25 @@ options(timeout=1000) #| message: false print("Loading Runsheet...") # NON_DPPD +df_rs <- read.csv(runsheet, check.names = FALSE) %>% + dplyr::mutate_all(function(x) iconv(x, "latin1", "ASCII", sub="")) # Convert all characters to ascii, when not possible, remove the character + +# NON_DPPD:START +print("Here is the embedded runsheet") +DT::datatable(df_rs) +# NON_DPPD:END +print("Loading Raw Data...") # NON_DPPD +local_paths <- ifelse( + stringr::str_detect(df_rs$`Array Data File Path`, "\\.gz$"), + stringr::str_remove(df_rs$`Array Data File Name`, "\\.gz$"), + df_rs$`Array Data File Name` +) +df_local_paths <- data.frame(`Sample Name` = df_rs$`Sample Name`, `Local Paths` = local_paths, check.names = FALSE) +# NON_DPPD:START +print("Raw Data Location Resolved") +DT::datatable(df_local_paths) +# NON_DPPD:END + # Utility function to improve robustness of function calls # Used to remedy intermittent internet issues during runtime retry_with_delay <- function(func, ...) { @@ -105,84 +129,13 @@ retry_with_delay <- function(func, ...) { } } -df_rs <- read.csv(runsheet, check.names = FALSE) %>% - dplyr::mutate_all(function(x) iconv(x, "latin1", "ASCII", sub="")) # Convert all characters to ascii, when not possible, remove the character - -# NON_DPPD:START -print("Here is the embedded runsheet") -DT::datatable(df_rs) -print("Here are the expected comparison groups") -# NON_DPPD:END -print("Loading Raw Data...") # NON_DPPD -allTrue <- function(i_vector) { - if ( length(i_vector) == 0 ) { - stop(paste("Input vector is length zero")) - } - all(i_vector) -} - -# Define paths to raw data files -runsheetPathsAreURIs <- function(df_runsheet) { - allTrue(stringr::str_starts(df_runsheet$`Array Data File Path`, "https")) -} - - -# Download raw data files -downloadFilesFromRunsheet <- function(df_runsheet) { - urls <- df_runsheet$`Array Data File Path` - destinationFiles <- df_runsheet$`Array Data File Name` - - mapply(function(url, destinationFile) { - print(paste0("Downloading from '", url, "' TO '", destinationFile, "'")) - if ( file.exists(destinationFile ) ) { - warning(paste( "Using Existing File:", destinationFile )) - } else { - download.file(url, destinationFile) - } - }, urls, destinationFiles) - - destinationFiles # Return these paths -} - - -if ( runsheetPathsAreURIs(df_rs) ) { - print("Determined Raw Data Locations are URIS") - local_paths <- retry_with_delay(downloadFilesFromRunsheet, df_rs) -} else { - print("Or Determined Raw Data Locations are local paths") - local_paths <- df_rs$`Array Data File Path` -} - - -# uncompress files if needed -if ( allTrue(stringr::str_ends(local_paths, ".gz")) ) { - print("Determined these files are gzip compressed... uncompressing now") - # This does the uncompression - lapply(local_paths, R.utils::gunzip, remove = FALSE, overwrite = TRUE) - # This removes the .gz extension to get the uncompressed filenames - local_paths <- vapply(local_paths, - stringr::str_replace, # Run this function against each item in 'local_paths' - FUN.VALUE = character(1), # Execpt an character vector as a return - USE.NAMES = FALSE, # Don't use the input to assign names for the returned list - pattern = ".gz$", # first argument for applied function - replacement = "" # second argument for applied function - ) -} - -df_local_paths <- data.frame(`Sample Name` = df_rs$`Sample Name`, `Local Paths` = local_paths, check.names = FALSE) -# NON_DPPD:START -print("Raw Data Loaded Successfully") -DT::datatable(df_local_paths) -# NON_DPPD:END - - # Load raw data into R object # Retry with delay here to accomodate oligo's automatic loading of annotation packages and occasional internet related failures to load raw_data <- retry_with_delay( oligo::read.celfiles, - df_local_paths$`Local Paths`, - sampleNames = df_local_paths$`Sample Name`# Map column names as Sample Names (instead of default filenames) - ) + df_local_paths$`Local Paths`, + sampleNames = df_local_paths$`Sample Name`# Map column names as Sample Names (instead of default filenames) + ) print(str(raw_data)) @@ -196,6 +149,9 @@ message(paste0("Number of Probes: ", dim(raw_data)[1])) # NON_DPPD DT::datatable(raw_data$targets, caption = "Sample to File Mapping") DT::datatable(head(raw_data$genes, n = 20), caption = "First 20 rows of raw data file embedded probes to genes table") # NON_DPPD:END + +annotation_file_path <- params$annotation_file_path +ensembl_version <- params$ensembl_version ``` ## QA For Raw Data @@ -461,9 +417,9 @@ par(original_par) #| message: false # Call RMA but skip normalize and background correction since those have already been applied probeset_level_data <- oligo::rma(norm_data, - normalize=FALSE, - background=FALSE, - ) + normalize=FALSE, + background=FALSE + ) # Summarize background-corrected and normalized data print("Summarized Probeset Level Data Below") # NON_DPPD @@ -479,15 +435,13 @@ DT::datatable(head(raw_data$genes, n = 20), caption = "First 20 rows of raw data # NON_DPPD:END ``` -## Perform Probeset Differential Expression and Annotation +## Probeset Annotations -### Probeset Differential Expression (DE) - -#### Add Probeset Annotations +### Get Probeset Annotations ``` {r retrieve-probeset-annotations} #| message: false -shortenedOrganismName <- function(long_name) { +shortened_organism_name <- function(long_name) { #' Convert organism names like 'Homo Sapiens' into 'hsapiens' tokens <- long_name %>% stringr::str_split(" ", simplify = TRUE) genus_name <- tokens[1] @@ -499,7 +453,7 @@ shortenedOrganismName <- function(long_name) { return(short_name) } -getBioMartAttribute <- function(df_rs) { +get_biomart_attribute <- function(df_rs) { #' Returns resolved biomart attribute source from runsheet # NON_DPPD:START #' this either comes from the runsheet or as a fall back, the parameters injected during render @@ -519,111 +473,161 @@ getBioMartAttribute <- function(df_rs) { } } -get_ensembl_genomes_mappings_from_ftp <- function(organism, ensembl_genomes_portal, ensembl_genomes_version, biomart_attribute) { - #' Obtain mapping table directly from ftp. Useful when biomart live service no longer exists for desired version - - request_url <- glue::glue("https://ftp.ebi.ac.uk/ensemblgenomes/pub/{ensembl_genomes_portal}/release-{ensembl_genomes_version}/mysql/{ensembl_genomes_portal}_mart_{ensembl_genomes_version}/{organism}_eg_gene__efg_{biomart_attribute}__dm.txt.gz") - - print(glue::glue("Mappings file URL: {request_url}")) - - # Create a temporary file name - temp_file <- tempfile(fileext = ".gz") +# Convert list of multi-mapped genes to string +list_to_unique_piped_string <- function(str_list) { + #! convert lists into strings denoting unique elements separated by '|' characters + #! e.g. c("GO1","GO2","GO2","G03") -> "GO1|GO2|GO3" + return(toString(unique(str_list)) %>% stringr::str_replace_all(pattern = stringr::fixed(", "), replacement = "|")) +} - # Download the gzipped table file using the download.file function - download.file(url = request_url, destfile = temp_file, method = "libcurl") # Use 'libcurl' to support ftps +# ---- Generic Ensembl FTP mart-dump helpers -------------------------------- +# Replaces live biomaRt::getBM() data queries with direct downloads of the same +# flat "mart dump" tables biomaRt itself queries under the hood, for both main +# Ensembl and Ensembl Genomes (plants/fungi/metazoa/protists/bacteria) releases. +# No retry/chunking/Sys.sleep: these are static per-release files on a stable +# FTP mirror, not a rate-limited live API, so a failed download is a real +# failure worth surfacing immediately rather than silently retrying. + +resolve_mart_ftp_base <- function(division, organism, ensembl_version, ensembl_genomes_portal = NULL) { + #' Resolve the FTP mart directory + per-species dataset prefix for either + #' the main Ensembl release ("main") or an Ensembl Genomes division ("genomes") + + if (division == "genomes") { + list( + ftp_dir = glue::glue("https://ftp.ebi.ac.uk/ensemblgenomes/pub/{ensembl_genomes_portal}/release-{ensembl_version}/mysql/{ensembl_genomes_portal}_mart_{ensembl_version}"), + dataset_prefix = glue::glue("{organism}_eg_gene") + ) + } else { + list( + ftp_dir = glue::glue("https://ftp.ebi.ac.uk/pub/ensembl/release-{ensembl_version}/mysql/ensembl_mart_{ensembl_version}"), + dataset_prefix = glue::glue("{organism}_gene_ensembl") + ) + } +} - # Uncompress the file - uncompressed_temp_file <- tempfile() - gzcon <- gzfile(temp_file, "rt") - content <- readLines(gzcon) - writeLines(content, uncompressed_temp_file) - close(gzcon) +download_mart_dump <- function(url, col_names = NULL) { + #' Download and load a single gzipped, tab-separated, headerless Ensembl "mart dump" table + options(timeout = 300) # Can be further increased for downloading large files + print(glue::glue("Mart dump URL: {url}")) - # Load the data into a dataframe - mapping <- read.table(uncompressed_temp_file, # Read the uncompressed file - # Add column names as follows: MAPID, TAIR, PROBESETID - col.names = c("MAPID", "ensembl_gene_id", biomart_attribute), - header = FALSE, # No header in original table - sep = "\t") # Tab separated + temp_file <- tempfile(fileext = ".gz") + download.file(url = url, destfile = temp_file, method = "libcurl") # libcurl needed for ftp(s) URLs - # Clean up temporary files + mapping <- read.table(gzfile(temp_file), header = FALSE, sep = "\t", quote = "", comment.char = "") + if (!is.null(col_names)) colnames(mapping)[seq_along(col_names)] <- col_names unlink(temp_file) - unlink(uncompressed_temp_file) - return(mapping) } +get_transcript_to_gene_mapping_from_ftp <- function(division, organism, ensembl_version, ensembl_genomes_portal = NULL) { + #' Transcript -> Gene, from the "*_transcript__main.txt.gz" mart dump. + #' Only used by the main-Ensembl probe->gene path (Probe -> Transcript -> Gene); + #' Ensembl Genomes divisions map probe -> gene directly and never call this. + #' Detects IDs by content instead of hardcoding column positions + + loc <- resolve_mart_ftp_base(division, organism, ensembl_version, ensembl_genomes_portal) + url <- glue::glue("{loc$ftp_dir}/{loc$dataset_prefix}__transcript__main.txt.gz") + raw <- retry_with_delay(download_mart_dump, url) # <- retried: large, always-expected file + + # Unversioned Ensembl gene/transcript IDs, e.g. ENSG00000210049 / ENSMUSG00000064336, + # ENST00000387314 / ENSMUST00000082387. Anchored with no trailing ".N" to + # exclude the separate versioned-ID columns present later in the same row. + gene_col_idx <- which(sapply(raw, function(col) any(grepl("^ENS[A-Z]*G[0-9]+$", col)))) + transcript_col_idx <- which(sapply(raw, function(col) any(grepl("^ENS[A-Z]*T[0-9]+$", col)))) + + stopifnot( + "Could not uniquely identify gene ID column in transcript main table" = length(gene_col_idx) == 1, + "Could not uniquely identify transcript ID column in transcript main table" = length(transcript_col_idx) == 1 + ) -organism <- shortenedOrganismName(unique(df_rs$organism)) - -if (organism %in% c('ecoli', 'paeruginosa')) { - expected_attribute_name <- 'ProbesetID' + raw %>% + dplyr::transmute( + ensembl_transcript_id = .data[[paste0("V", transcript_col_idx)]], + ensembl_gene_id = .data[[paste0("V", gene_col_idx)]] + ) %>% + dplyr::distinct() +} - annot_file <- c( - 'ecoli' = 'E_coli_2.na36.annot.csv', - 'paeruginosa' = 'Pae_G1a.na36.annot.csv' +get_probe_to_gene_mapping_from_ftp <- function(division, organism, ensembl_version, biomart_attribute, ensembl_genomes_portal = NULL) { + #' Probe -> Gene. The two BioMart divisions use different schemas here: + #' Main Ensembl ("main"): the efg dm file maps probe -> TRANSCRIPT; a + #' second join against the transcript__main table is required to reach gene. + #' Ensembl Genomes ("genomes", e.g. Ensembl Plants): the efg dm file maps + #' probe -> GENE directly for microarray platforms. No second join exists + #' here -- attempting one would silently produce wrong/empty results, + #' since the dm file's second column is already a gene ID, not a transcript ID. + #' Only the attribute-specific dm file (probe_dm below) is + #' expected to legitimately be missing -- not every array design has a + #' registered efg_* attribute in Ensembl BioMart. A failure there is + #' treated as "fall back to custom annotation" (returns NULL), not an + #' error. transcript__main.txt.gz, in contrast, is a large stable file + #' that should always exist for a given release; a failure there is a + #' real problem (network/transient) and is allowed to propagate as an + #' actual error rather than being silently swallowed into a fallback. + + loc <- resolve_mart_ftp_base(division, organism, ensembl_version, ensembl_genomes_portal) + dm_url <- glue::glue("{loc$ftp_dir}/{loc$dataset_prefix}__efg_{biomart_attribute}__dm.txt.gz") + + probe_dm <- tryCatch( + if (division == "genomes") { + download_mart_dump(dm_url, col_names = c("MAPID", "ensembl_gene_id", biomart_attribute)) + } else { + download_mart_dump(dm_url, col_names = c("MAPID", "ensembl_transcript_id", biomart_attribute)) + }, + error = function(e) { + message(glue::glue("Probe mapping file not available for attribute '{biomart_attribute}' ({e$message}); falling back to custom annotation")) + NULL + } ) - df_mapping <- read.csv( - file.path(params$local_annotation_dir, annot_file[[organism]]), - skip = 13, header = TRUE, na.strings = c('NA', '---') - )[c('Probe.Set.ID', 'Entrez.Gene', 'Gene.Symbol', 'Gene.Title', 'Ensembl', 'RefSeq.Transcript.ID', 'RefSeq.Protein.ID', 'Gene.Ontology.Biological.Process', 'Gene.Ontology.Cellular.Component', 'Gene.Ontology.Molecular.Function')] + if (is.null(probe_dm)) return(NULL) - # Clean columns - df_mapping$Gene.Symbol <- purrr::map_chr(stringr::str_split(df_mapping$Gene.Symbol, stringr::fixed(' /// ')), ~paste0(unique(.), collapse = "|")) %>% stringr::str_replace('NA', NA_character_) - df_mapping$Gene.Title <- purrr::map_chr(stringr::str_split(df_mapping$Gene.Title, stringr::fixed(' /// ')), ~paste0(unique(.), collapse = "|")) %>% stringr::str_replace('NA', NA_character_) - df_mapping$Entrez.Gene <- purrr::map_chr(stringr::str_split(df_mapping$Entrez.Gene, stringr::fixed(' /// ')), ~paste0(unique(.), collapse = "|")) %>% stringr::str_replace('NA', NA_character_) - df_mapping$Ensembl <- purrr::map_chr(stringr::str_split(df_mapping$Ensembl, stringr::fixed(' /// ')), ~paste0(unique(.), collapse = "|")) %>% stringr::str_replace('NA', NA_character_) + if (division == "genomes") { + return(probe_dm %>% dplyr::select(!!sym(biomart_attribute), ensembl_gene_id)) + } - df_mapping$RefSeq <- paste(df_mapping$RefSeq.Transcript.ID, df_mapping$RefSeq.Protein.ID) - df_mapping$RefSeq <- purrr::map_chr(stringr::str_extract_all(df_mapping$RefSeq, '[A-Z]+_[\\d.]+'), ~paste0(unique(.), collapse = "|")) %>% stringr::str_replace('^$', NA_character_) + transcript_to_gene <- get_transcript_to_gene_mapping_from_ftp(division, organism, ensembl_version, ensembl_genomes_portal) - df_mapping$GO <- paste(df_mapping$Gene.Ontology.Biological.Process, df_mapping$Gene.Ontology.Cellular.Component, df_mapping$Gene.Ontology.Molecular.Function) - df_mapping$GO <- purrr::map_chr(stringr::str_extract_all(df_mapping$GO, '\\d{7}'), ~paste0('GO:', unique(.), collapse = "|")) %>% stringr::str_replace('^GO:$', NA_character_) + probe_dm %>% + dplyr::left_join(transcript_to_gene, by = "ensembl_transcript_id") %>% + dplyr::select(!!sym(biomart_attribute), ensembl_gene_id) +} - df_mapping <- df_mapping[c('Probe.Set.ID', 'Entrez.Gene', 'Gene.Symbol', 'Gene.Title', 'Ensembl', 'RefSeq', 'GO')] - names(df_mapping) <- c('ProbesetID', 'ENTREZID', 'SYMBOL', 'GENENAME', 'ENSEMBL', 'REFSEQ', 'GOSLIM_IDS') +# ---- Main probe annotation resolution ------------------------------------- +organism <- shortened_organism_name(unique(df_rs$organism)) +annot_key <- ifelse(organism %in% c("athaliana"), 'TAIR', 'ENSEMBL') +ENSEMBL_VERSION <- ensembl_version +expected_attribute_name <- get_biomart_attribute(df_rs) - df_mapping$STRING_id <- NA_character_ -} else if (organism %in% c("athaliana")) { - ensembl_genomes_version = params$ensembl_version +if (organism %in% c("athaliana")) { + # Unchanged from the prior implementation: single-step Probe -> Gene lookup + # via Ensembl Genomes' own efg dm file, whose second column is already a + # gene ID for this division (not a transcript ID needing a further join). + ensembl_genomes_portal = "plants" - print(glue::glue("Using ensembl genomes ftp to get specific version of probeset id mapping table. Ensembl genomes portal: {ensembl_genomes_portal}, version: {ensembl_genomes_version}")) - expected_attribute_name <- getBioMartAttribute(df_rs) - df_mapping <- retry_with_delay( - get_ensembl_genomes_mappings_from_ftp, + print(glue::glue("Using Ensembl Genomes FTP to get probe mapping table. Portal: {ensembl_genomes_portal}, version: {ENSEMBL_VERSION}")) + + df_mapping <- get_probe_to_gene_mapping_from_ftp( + division = "genomes", organism = organism, - ensembl_genomes_portal = ensembl_genomes_portal, - ensembl_genomes_version = ensembl_genomes_version, - biomart_attribute = expected_attribute_name + ensembl_version = ENSEMBL_VERSION, + biomart_attribute = expected_attribute_name, + ensembl_genomes_portal = ensembl_genomes_portal ) - # TAIR from the mapping tables tend to be in the format 'AT1G01010.1' but the raw data has 'AT1G01010' + # TAIR IDs in the mapping tables tend to be in the format 'AT1G01010.1' but the raw data has 'AT1G01010' # So here we remove the '.NNN' from the mapping table where .NNN is any number df_mapping$ensembl_gene_id <- stringr::str_replace_all(df_mapping$ensembl_gene_id, "\\.\\d+$", "") -} else { - # Use biomart from main Ensembl website which archives keep each release on the live service - # locate dataset - expected_dataset_name <- shortenedOrganismName(unique(df_rs$organism)) %>% stringr::str_c("_gene_ensembl") - print(paste0("Expected dataset name: '", expected_dataset_name, "'")) - message(paste0("Expected dataset name: '", expected_dataset_name, "'")) # NON_DPPD - - - # Specify Ensembl version used in current GeneLab reference annotations - ENSEMBL_VERSION <- params$ensembl_version - print(paste0("Searching for Ensembl Version: ", ENSEMBL_VERSION)) # NON_DPPD - print(glue::glue("Using Ensembl biomart to get specific version of mapping table. Ensembl version: {ENSEMBL_VERSION}")) - - ensembl <- biomaRt::useEnsembl(biomart = "genes", - dataset = expected_dataset_name, - version = ENSEMBL_VERSION) - print(ensembl) - - expected_attribute_name <- getBioMartAttribute(df_rs) - print(paste0("Expected attribute name: '", expected_attribute_name, "'")) - message(paste0("Expected attribute name: '", expected_attribute_name, "'")) # NON_DPPD + use_custom_annot <- FALSE +} else { + expected_dataset_name <- glue::glue("{organism}_gene_ensembl") + print(glue::glue("Expected dataset name: '{expected_dataset_name}'")) + message(glue::glue("Expected dataset name: '{expected_dataset_name}'")) # NON_DPPD + print(glue::glue("Expected attribute name: '{expected_attribute_name}'")) + message(glue::glue("Expected attribute name: '{expected_attribute_name}'")) # NON_DPPD + print(glue::glue("Searching for Ensembl Version: {ENSEMBL_VERSION}")) # NON_DPPD # Some probe_ids for affy_hta_2_0 may end in .hg.1 instead of .hg (how it is in biomaRt), leading to 0 results returned if (expected_attribute_name == 'affy_hta_2_0') { @@ -632,73 +636,93 @@ if (organism %in% c('ecoli', 'paeruginosa')) { probe_ids <- rownames(probeset_level_data) - # DEBUG:START - if ( is.integer(params$DEBUG_limit_biomart_query) ) { - warning(paste("DEBUG MODE: Limiting query to", params$DEBUG_limit_biomart_query, "entries")) - message(paste("DEBUG MODE: Limiting query to", params$DEBUG_limit_biomart_query, "entries")) - probe_ids <- probe_ids[1:params$DEBUG_limit_biomart_query] + print(glue::glue("Using Ensembl biomart to get specific version of mapping table. Ensembl version: {ENSEMBL_VERSION}")) + print(glue::glue("Attempting Ensembl FTP probe mapping table. Version: {ENSEMBL_VERSION}")) + df_mapping <- get_probe_to_gene_mapping_from_ftp( + division = "main", + organism = organism, + ensembl_version = ENSEMBL_VERSION, + biomart_attribute = expected_attribute_name + ) + + if (!is.null(df_mapping)) { + use_custom_annot <- FALSE + # FTP has no server-side filter equivalent to getBM(filters=, values=) + # the full mart-dump file is always downloaded in full, then filtered + # client-side down to just this experiment's probes. + df_mapping <- df_mapping %>% dplyr::filter(!!sym(expected_attribute_name) %in% probe_ids) + } else { + use_custom_annot <- TRUE } - # DEBUG:END - - # Create probe map - # Run Biomart Queries in chunks to prevent request timeouts - # Note: If timeout is occuring (possibly due to larger load on biomart), reduce chunk size - CHUNK_SIZE= 1500 - probe_id_chunks <- split(probe_ids, ceiling(seq_along(probe_ids) / CHUNK_SIZE)) - df_mapping <- data.frame() - for (i in seq_along(probe_id_chunks)) { - probe_id_chunk <- probe_id_chunks[[i]] - print(glue::glue("Running biomart query chunk {i} of {length(probe_id_chunks)}. Total probes IDS in query ({length(probe_id_chunk)})")) - message(glue::glue("Running biomart query chunk {i} of {length(probe_id_chunks)}. Total probes IDS in query ({length(probe_id_chunk)})")) # NON_DPPD - chunk_results <- biomaRt::getBM( - attributes = c( - expected_attribute_name, - "ensembl_gene_id" - ), - filters = expected_attribute_name, - values = probe_id_chunk, - mart = ensembl) - - if (nrow(chunk_results) > 0) { - df_mapping <- df_mapping %>% dplyr::bind_rows(chunk_results) +} + +# At this point, we have df_mapping from either the Ensembl FTP mart dumps (main or Ensembl Genomes) depending on the organism +# If no df_mapping obtained (e.g., organism not supported in biomart), use custom annotations; otherwise, merge in-house annotations to df_mapping + +if (use_custom_annot) { + expected_attribute_name <- 'ProbesetID' + annot_type <- 'NO_CUSTOM_ANNOT' + + if (!is.null(local_annotation_dir) && !is.null(array_annot_path)) { + probe_annot_df <- read.csv(array_annot_path, row.names=1) + if (unique(df_rs$`biomart_attribute`) %in% row.names(probe_annot_df)) { + annot_config <- probe_annot_df[unique(df_rs$`biomart_attribute`), ] + annot_type <- annot_config$annot_type[[1]] + } else { + warning(paste0("No entry for '", unique(df_rs$`biomart_attribute`), "' in provided custom probe annotation file: ", array_annot_path)) } - - Sys.sleep(10) # Slight break between requests to prevent back-to-back requests + } else { + warning("Need to provide both local_annotation_dir and array_annot_path to use custom annotation.") } -} -# At this point, we have df_mapping from either the biomart live service or the ensembl genomes ftp archive depending on the organism -``` + if (annot_type == '3prime-IVT') { + unique_probe_ids <- read.csv( + file.path(local_annotation_dir, annot_config$annot_filename[[1]]), + skip = 13, header = TRUE, na.strings = c('NA', '---') + )[c('Probe.Set.ID', 'Entrez.Gene', 'Gene.Symbol', 'Gene.Title', 'Ensembl', 'RefSeq.Transcript.ID', 'RefSeq.Protein.ID', 'Gene.Ontology.Biological.Process', 'Gene.Ontology.Cellular.Component', 'Gene.Ontology.Molecular.Function')] -``` {r reformat-merge-probe-annotations} -# Convert list of multi-mapped genes to string -listToUniquePipedString <- function(str_list) { - #! convert lists into strings denoting unique elements separated by '|' characters - #! e.g. c("GO1","GO2","GO2","G03") -> "GO1|GO2|GO3" - return(toString(unique(str_list)) %>% stringr::str_replace_all(pattern = stringr::fixed(", "), replacement = "|")) -} + # Clean columns + unique_probe_ids$Gene.Symbol <- purrr::map_chr(stringr::str_split(unique_probe_ids$Gene.Symbol, stringr::fixed(' /// ')), ~paste0(unique(.), collapse = "|")) %>% stringr::str_replace('NA', NA_character_) + unique_probe_ids$Gene.Title <- purrr::map_chr(stringr::str_split(unique_probe_ids$Gene.Title, stringr::fixed(' /// ')), ~paste0(unique(.), collapse = "|")) %>% stringr::str_replace('NA', NA_character_) + unique_probe_ids$Entrez.Gene <- purrr::map_chr(stringr::str_split(unique_probe_ids$Entrez.Gene, stringr::fixed(' /// ')), ~paste0(unique(.), collapse = "|")) %>% stringr::str_replace('NA', NA_character_) + unique_probe_ids$Ensembl <- purrr::map_chr(stringr::str_split(unique_probe_ids$Ensembl, stringr::fixed(' /// ')), ~paste0(unique(.), collapse = "|")) %>% stringr::str_replace('NA', NA_character_) -annot_key = ifelse(organism %in% c("athaliana"), 'TAIR', 'ENSEMBL') + unique_probe_ids$RefSeq <- paste(unique_probe_ids$RefSeq.Transcript.ID, unique_probe_ids$RefSeq.Protein.ID) + unique_probe_ids$RefSeq <- purrr::map_chr(stringr::str_extract_all(unique_probe_ids$RefSeq, '[A-Z]+_[\\d.]+'), ~paste0(unique(.), collapse = "|")) %>% stringr::str_replace('^$', NA_character_) -if (organism %in% c('ecoli')) { - unique_probe_ids <- df_mapping %>% - dplyr::mutate( - count_ENTREZID_mappings = 1 + stringr::str_count(ENTREZID, stringr::fixed("|")) - ) + unique_probe_ids$GO <- paste(unique_probe_ids$Gene.Ontology.Biological.Process, unique_probe_ids$Gene.Ontology.Cellular.Component, unique_probe_ids$Gene.Ontology.Molecular.Function) + unique_probe_ids$GO <- purrr::map_chr(stringr::str_extract_all(unique_probe_ids$GO, '\\d{7}'), ~paste0('GO:', unique(.), collapse = "|")) %>% stringr::str_replace('^GO:$', NA_character_) + + unique_probe_ids <- unique_probe_ids[c('Probe.Set.ID', 'Entrez.Gene', 'Gene.Symbol', 'Gene.Title', 'Ensembl', 'RefSeq', 'GO')] + names(unique_probe_ids) <- c('ProbesetID', 'ENTREZID', 'SYMBOL', 'GENENAME', 'ENSEMBL', 'REFSEQ', 'GOSLIM_IDS') - primary_key <- 'ENTREZID' - primary_key_count <- 'count_ENTREZID_mappings' -} else if (organism %in% c('paeruginosa')) { - unique_probe_ids <- df_mapping %>% + unique_probe_ids$STRING_id <- NA_character_ + + gene_col <- 'ENSEMBL' + if (sum(!is.na(unique_probe_ids$ENTREZID)) > sum(!is.na(unique_probe_ids$ENSEMBL))) { + gene_col <- 'ENTREZID' + } + if (sum(!is.na(unique_probe_ids$SYMBOL)) > max(sum(!is.na(unique_probe_ids$ENTREZID)), sum(!is.na(unique_probe_ids$ENSEMBL)))) { + gene_col <- 'SYMBOL' + } + + unique_probe_ids <- unique_probe_ids %>% dplyr::mutate( - count_SYMBOL_mappings = 1 + stringr::str_count(SYMBOL, stringr::fixed("|")) + count_gene_mappings = 1 + stringr::str_count(get(gene_col), stringr::fixed("|")), + gene_mapping_source = gene_col ) - - primary_key <- 'SYMBOL' - primary_key_count <- 'count_SYMBOL_mappings' + } else if (annot_type == 'custom') { + unique_probe_ids <- read.csv( + file.path(local_annotation_dir, annot_config$annot_filename[[1]]), + header = TRUE, na.strings = c('NA', '') + ) + } else { + annot_cols <- c('ProbesetID', 'ENTREZID', 'SYMBOL', 'GENENAME', 'ENSEMBL', 'REFSEQ', 'GOSLIM_IDS', 'STRING_id', 'count_gene_mappings', 'gene_mapping_source') + unique_probe_ids <- setNames(data.frame(matrix(NA_character_, nrow = 1, ncol = length(annot_cols))), annot_cols) + } } else { annot <- read.table( - as.character(params$annotation_file_path), + as.character(annotation_file_path), sep = "\t", header = TRUE, quote = "", @@ -710,37 +734,37 @@ if (organism %in% c('ecoli')) { dplyr::mutate(dplyr::across(!!sym(expected_attribute_name), as.character)) %>% # Ensure probeset ids treated as character type dplyr::group_by(!!sym(expected_attribute_name)) %>% dplyr::summarise( - ENSEMBL = listToUniquePipedString(ensembl_gene_id) + ENSEMBL = list_to_unique_piped_string(ensembl_gene_id) ) %>% # Count number of ensembl IDS mapped dplyr::mutate( - count_ENSEMBL_mappings = 1 + stringr::str_count(ENSEMBL, stringr::fixed("|")) + count_gene_mappings = 1 + stringr::str_count(ENSEMBL, stringr::fixed("|")), + gene_mapping_source = annot_key ) %>% dplyr::left_join(annot, by = c("ENSEMBL" = annot_key)) - - - primary_key <- 'ENSEMBL' - primary_key_count <- 'count_ENSEMBL_mappings' } +``` +``` {r reformat-merge-probe-annotations} probeset_expression_matrix <- oligo::exprs(probeset_level_data) -probeset_expression_matrix.biomart_mapped <- probeset_expression_matrix %>% +probeset_expression_matrix.gene_mapped <- probeset_expression_matrix %>% as.data.frame() %>% tibble::rownames_to_column(var = "ProbesetID") %>% # Ensure rownames (probeset IDs) can be used as join key dplyr::left_join(unique_probe_ids, by = c("ProbesetID" = expected_attribute_name ) ) %>% - dplyr::mutate( !!primary_key_count := ifelse(is.na(get(primary_key)), 0, get(primary_key_count)) ) + dplyr::mutate( count_gene_mappings := ifelse(is.na(count_gene_mappings), 0, count_gene_mappings) ) %>% + dplyr::mutate( gene_mapping_source := unique(unique_probe_ids$gene_mapping_source) ) ``` -### Summarize Probeset Mapping +### Summarize Gene Mapping ``` {r summarize-remapping-vs-original-mapping} #| message: false # Pie Chart with Percentages slices <- c( - 'Unique Mapping' = nrow(probeset_expression_matrix.biomart_mapped %>% dplyr::filter(get(primary_key_count) == 1) %>% dplyr::distinct(ProbesetID)), - 'Multi Mapping' = nrow(probeset_expression_matrix.biomart_mapped %>% dplyr::filter(get(primary_key_count) > 1) %>% dplyr::distinct(ProbesetID)), - 'No Mapping' = nrow(probeset_expression_matrix.biomart_mapped %>% dplyr::filter(get(primary_key_count) == 0) %>% dplyr::distinct(ProbesetID)) + 'Unique Mapping' = nrow(probeset_expression_matrix.gene_mapped %>% dplyr::filter(count_gene_mappings == 1) %>% dplyr::distinct(ProbesetID)), + 'Multi Mapping' = nrow(probeset_expression_matrix.gene_mapped %>% dplyr::filter(count_gene_mappings > 1) %>% dplyr::distinct(ProbesetID)), + 'No Mapping' = nrow(probeset_expression_matrix.gene_mapped %>% dplyr::filter(count_gene_mappings == 0) %>% dplyr::distinct(ProbesetID)) ) pct <- round(slices/sum(slices)*100) chart_names <- names(slices) @@ -748,17 +772,104 @@ chart_names <- glue::glue("{names(slices)} ({slices})") # add count to labels chart_names <- paste(chart_names, pct) # add percents to labels chart_names <- paste(chart_names,"%",sep="") # ad % to labels pie(slices,labels = chart_names, col=rainbow(length(slices)), - main=glue::glue("Mapping to Primary Keytype\n {nrow(probeset_expression_matrix.biomart_mapped %>% dplyr::distinct(ProbesetID))} Total Unique Probesets") + main=glue::glue("Mapping to Primary Keytype\n {nrow(probeset_expression_matrix.gene_mapped %>% dplyr::distinct(ProbesetID))} Total Unique Probesets") ) print(glue::glue("Unique Mapping Count: {slices[['Unique Mapping']]}")) message(glue::glue("Unique Mapping Count: {slices[['Unique Mapping']]}")) # NON_DPPD ``` +### Generate Annotated Raw and Normalized Expression Tables + +```{r save-tables} +## Reorder columns before saving to file +ANNOTATIONS_COLUMN_ORDER = c( + annot_key, + "SYMBOL", + "GENENAME", + "REFSEQ", + "ENTREZID", + "STRING_id", + "GOSLIM_IDS" +) + +SAMPLE_COLUMN_ORDER <- df_rs$`Sample Name` + +probeset_expression_matrix.gene_mapped <- probeset_expression_matrix.gene_mapped %>% dplyr::rename( !!annot_key := ENSEMBL ) + +## Output column subset file with just normalized probeset level expression values +write.csv( + probeset_expression_matrix.gene_mapped[c( + ANNOTATIONS_COLUMN_ORDER, + "ProbesetID", + "count_gene_mappings", + "gene_mapping_source", + SAMPLE_COLUMN_ORDER) + ], file.path(DIR_NORMALIZED_EXPRESSION, "normalized_expression_probeset_GLmicroarray.csv"), row.names = FALSE) + +## Determine column order for probe level tables + +PROBE_INFO_COLUMN_ORDER = c( + "ProbesetID", + "ProbeID", + "count_gene_mappings", + "gene_mapping_source" +) + +FINAL_COLUMN_ORDER <- c( + ANNOTATIONS_COLUMN_ORDER, + PROBE_INFO_COLUMN_ORDER, + SAMPLE_COLUMN_ORDER +) + +## Generate raw intensity matrix that includes annotations + +background_corrected_data_annotated <- oligo::exprs(background_corrected_data) %>% + as.data.frame() %>% + tibble::rownames_to_column(var = "fid") %>% # Ensure rownames (probeset IDs) can be used as join key + dplyr::mutate(dplyr::across(fid, as.integer)) %>% # Ensure fid is integer type, consistent with getProbeInfo typing + dplyr::right_join(oligo::getProbeInfo(background_corrected_data), by = "fid") %>% # Add 'man_fsetid' via mapping based on fid + dplyr::rename( ProbesetID = man_fsetid ) %>% # Rename from getProbeInfo name to ProbesetID + dplyr::rename( ProbeID = fid ) %>% # Rename from getProbeInfo name to ProbeID + dplyr::left_join(unique_probe_ids, by = c("ProbesetID" = expected_attribute_name ) ) %>% # Join with biomaRt ENSEMBL mappings + dplyr::mutate( count_gene_mappings := ifelse(is.na(count_gene_mappings), 0, count_gene_mappings) ) %>% # Convert NA mapping to 0 + dplyr::mutate( gene_mapping_source := unique(unique_probe_ids$gene_mapping_source) ) %>% + dplyr::rename( !!annot_key := ENSEMBL ) + +## Perform reordering +background_corrected_data_annotated <- background_corrected_data_annotated %>% + dplyr::relocate(dplyr::all_of(FINAL_COLUMN_ORDER)) + +write.csv(background_corrected_data_annotated, file.path(DIR_RAW_DATA, "raw_intensities_probe_GLmicroarray.csv"), row.names = FALSE) + +## Generate normalized expression matrix that includes annotations +norm_data_matrix_annotated <- oligo::exprs(norm_data) %>% + as.data.frame() %>% + tibble::rownames_to_column(var = "fid") %>% # Ensure rownames (probeset IDs) can be used as join key + dplyr::mutate(dplyr::across(fid, as.integer)) %>% # Ensure fid is integer type, consistent with getProbeInfo typing + dplyr::right_join(oligo::getProbeInfo(norm_data), by = "fid") %>% # Add 'man_fsetid' via mapping based on fid + dplyr::rename( ProbesetID = man_fsetid ) %>% # Rename from getProbeInfo name to ProbesetID + dplyr::rename( ProbeID = fid ) %>% # Rename from getProbeInfo name to ProbeID + dplyr::left_join(unique_probe_ids, by = c("ProbesetID" = expected_attribute_name ) ) %>% + dplyr::mutate( count_gene_mappings := ifelse(is.na(count_gene_mappings), 0, count_gene_mappings) ) %>% # Convert NA mapping to 0 + dplyr::mutate( gene_mapping_source := unique(unique_probe_ids$gene_mapping_source) ) %>% + dplyr::rename( !!annot_key := ENSEMBL ) + +norm_data_matrix_annotated <- norm_data_matrix_annotated %>% + dplyr::relocate(dplyr::all_of(FINAL_COLUMN_ORDER)) + +write.csv(norm_data_matrix_annotated, file.path(DIR_NORMALIZED_EXPRESSION, "normalized_intensities_probe_GLmicroarray.csv"), row.names = FALSE) +``` + +## Perform Probeset Differential Expression (DE) + ### Generate Design Matrix ``` {r generate-design-matrix} -runsheetToDesignMatrix <- function(runsheet_path) { +#| include: !expr params$run_DE +#| eval: !expr params$run_DE + +runsheet_to_design_matrix <- function(runsheet_path) { df <- read.csv(runsheet, check.names = FALSE) %>% dplyr::mutate_all(function(x) iconv(x, "latin1", "ASCII", sub="")) # Convert all characters to ascii, when not possible, remove the character # get only Factor Value columns factors = as.data.frame(df[,grep("Factor.Value", colnames(df), ignore.case=TRUE)]) @@ -801,14 +912,12 @@ runsheetToDesignMatrix <- function(runsheet_path) { # Loading metadata from runsheet csv file -design_data <- runsheetToDesignMatrix(runsheet) +design_data <- runsheet_to_design_matrix(runsheet) design <- design_data$matrix -if (params$run_DE) { - # Write SampleTable.csv and contrasts.csv file - write.csv(design_data$groups, file.path(DIR_DGE, "SampleTable_GLmicroarray.csv"), row.names = FALSE) - write.csv(design_data$contrasts, file.path(DIR_DGE, "contrasts_GLmicroarray.csv")) -} +# Write SampleTable.csv and contrasts.csv file +write.csv(design_data$groups, file.path(DIR_DGE, "SampleTable_GLmicroarray.csv"), row.names = FALSE) +write.csv(design_data$contrasts, file.path(DIR_DGE, "contrasts_GLmicroarray.csv")) ``` ### Perform Individual Probeset Level DE @@ -817,7 +926,7 @@ if (params$run_DE) { #| include: !expr params$run_DE #| eval: !expr params$run_DE -lmFitPairwise <- function(norm_data, design) { +lm_fit_pairwise <- function(norm_data, design) { #' Perform all pairwise comparisons #' Approach based on limma manual section 17.4 (version 3.52.4) @@ -836,7 +945,7 @@ lmFitPairwise <- function(norm_data, design) { } # Calculate results -res <- lmFitPairwise(probeset_level_data, design) +res <- lm_fit_pairwise(probeset_level_data, design) DT::datatable(limma::topTable(res)) # NON_DPPD # Print DE table, without filtering @@ -847,89 +956,7 @@ limma::write.fit(res, adjust = 'BH', sep = ",") ``` -### Add Additional Columns and Format DE Table - -```{r save-tables} -## Reorder columns before saving to file -ANNOTATIONS_COLUMN_ORDER = c( - annot_key, - "SYMBOL", - "GENENAME", - "REFSEQ", - "ENTREZID", - "STRING_id", - "GOSLIM_IDS" -) - -SAMPLE_COLUMN_ORDER <- design_data$group %>% dplyr::pull(sample) - -probeset_expression_matrix.biomart_mapped <- probeset_expression_matrix.biomart_mapped %>% dplyr::rename( !!annot_key := ENSEMBL ) - -## Output column subset file with just normalized probeset level expression values -write.csv( - probeset_expression_matrix.biomart_mapped[c( - ANNOTATIONS_COLUMN_ORDER, - "ProbesetID", - primary_key_count, - SAMPLE_COLUMN_ORDER) - ], file.path(DIR_NORMALIZED_EXPRESSION, "normalized_expression_probeset_GLmicroarray.csv"), row.names = FALSE) - -if (params$run_DE) { - ### Generate and export PCA table for GeneLab visualization plots - PCA_raw <- prcomp(t(exprs(probeset_level_data)), scale = FALSE) # Note: expression at the Probeset level is already log2 transformed - write.csv(PCA_raw$x, file.path(DIR_DGE, "visualization_PCA_table_GLmicroarray.csv")) -} - -## Determine column order for probe level tables - -PROBE_INFO_COLUMN_ORDER = c( - "ProbesetID", - "ProbeID", - primary_key_count -) - -FINAL_COLUMN_ORDER <- c( - ANNOTATIONS_COLUMN_ORDER, - PROBE_INFO_COLUMN_ORDER, - SAMPLE_COLUMN_ORDER - ) - -## Generate raw intensity matrix that includes annotations - -background_corrected_data_annotated <- oligo::exprs(background_corrected_data) %>% - as.data.frame() %>% - tibble::rownames_to_column(var = "fid") %>% # Ensure rownames (probeset IDs) can be used as join key - dplyr::mutate(dplyr::across(fid, as.integer)) %>% # Ensure fid is integer type, consistent with getProbeInfo typing - dplyr::right_join(oligo::getProbeInfo(background_corrected_data), by = "fid") %>% # Add 'man_fsetid' via mapping based on fid - dplyr::rename( ProbesetID = man_fsetid ) %>% # Rename from getProbeInfo name to ProbesetID - dplyr::rename( ProbeID = fid ) %>% # Rename from getProbeInfo name to ProbeID - dplyr::left_join(unique_probe_ids, by = c("ProbesetID" = expected_attribute_name ) ) %>% # Join with biomaRt ENSEMBL mappings - dplyr::mutate( !!primary_key_count := ifelse(is.na(get(primary_key)), 0, get(primary_key_count)) ) %>% # Convert NA mapping to 0 - dplyr::rename( !!annot_key := ENSEMBL ) - -## Perform reordering -background_corrected_data_annotated <- background_corrected_data_annotated %>% - dplyr::relocate(dplyr::all_of(FINAL_COLUMN_ORDER)) - -write.csv(background_corrected_data_annotated, file.path(DIR_RAW_DATA, "raw_intensities_probe_GLmicroarray.csv"), row.names = FALSE) - -## Generate normalized expression matrix that includes annotations -norm_data_matrix_annotated <- oligo::exprs(norm_data) %>% - as.data.frame() %>% - tibble::rownames_to_column(var = "fid") %>% # Ensure rownames (probeset IDs) can be used as join key - dplyr::mutate(dplyr::across(fid, as.integer)) %>% # Ensure fid is integer type, consistent with getProbeInfo typing - dplyr::right_join(oligo::getProbeInfo(norm_data), by = "fid") %>% # Add 'man_fsetid' via mapping based on fid - dplyr::rename( ProbesetID = man_fsetid ) %>% # Rename from getProbeInfo name to ProbesetID - dplyr::rename( ProbeID = fid ) %>% # Rename from getProbeInfo name to ProbeID - dplyr::left_join(unique_probe_ids, by = c("ProbesetID" = expected_attribute_name ) ) %>% - dplyr::mutate( !!primary_key_count := ifelse(is.na(get(primary_key)), 0, get(primary_key_count)) ) %>% # Convert NA mapping to 0 - dplyr::rename( !!annot_key := ENSEMBL ) - -norm_data_matrix_annotated <- norm_data_matrix_annotated %>% - dplyr::relocate(dplyr::all_of(FINAL_COLUMN_ORDER)) - -write.csv(norm_data_matrix_annotated, file.path(DIR_NORMALIZED_EXPRESSION, "normalized_intensities_probe_GLmicroarray.csv"), row.names = FALSE) -``` +### Add Annotation and Stats Columns and Format DE Table ``` {r save-de-table} #| message: false @@ -940,9 +967,9 @@ write.csv(norm_data_matrix_annotated, file.path(DIR_NORMALIZED_EXPRESSION, "norm # Read in DE table df_interim <- read.csv("INTERIM.csv") -# Bind columns from biomart mapped expression table +# Bind columns from gene mapped expression table df_interim <- df_interim %>% - dplyr::bind_cols(probeset_expression_matrix.biomart_mapped) + dplyr::bind_cols(probeset_expression_matrix.gene_mapped) # Reformat column names reformat_names <- function(colname, group_name_mapping) { @@ -985,15 +1012,10 @@ df_interim <- df_interim %>% dplyr::rename_with(reformat_names, .cols = matches( unique_groups <- unique(design_data$group$group) for ( i in seq_along(unique_groups) ) { current_group <- unique_groups[i] - current_samples <- design_data$group %>% - dplyr::group_by(group) %>% - dplyr::summarize( - samples = sort(unique(sample)) - ) %>% - dplyr::filter( - group == current_group - ) %>% - dplyr::pull() + current_samples <- design_data$group %>% + dplyr::filter(group == current_group) %>% + dplyr::pull(sample) %>% + sort() print(glue::glue("Computing mean and standard deviation for Group {i} of {length(unique_groups)}")) print(glue::glue("Group: {current_group}")) @@ -1036,7 +1058,8 @@ df_interim <- df_interim %>% dplyr::select(-any_of(colnames_to_remove)) PROBE_INFO_COLUMN_ORDER = c( "ProbesetID", - primary_key_count + "count_gene_mappings", + "gene_mapping_source" ) generate_prefixed_column_order <- function(subjects, prefixes) { @@ -1071,27 +1094,22 @@ ALL_SAMPLE_STATS_COLUMNS_ORDER <- c( "F.p.value" ) -GROUP_MEAN_COLUMNS_ORDER <- generate_prefixed_column_order( - subjects = unique(design_data$groups$group), - prefixes = c( - "Group.Mean_" - ) - ) -GROUP_STDEV_COLUMNS_ORDER <- generate_prefixed_column_order( +GROUP_MEAN_STDEV_COLUMNS_ORDER <- generate_prefixed_column_order( subjects = unique(design_data$groups$group), prefixes = c( + "Group.Mean_", "Group.Stdev_" - ) ) +) + FINAL_COLUMN_ORDER <- c( ANNOTATIONS_COLUMN_ORDER, PROBE_INFO_COLUMN_ORDER, SAMPLE_COLUMN_ORDER, STAT_COLUMNS_ORDER, ALL_SAMPLE_STATS_COLUMNS_ORDER, - GROUP_MEAN_COLUMNS_ORDER, - GROUP_STDEV_COLUMNS_ORDER - ) + GROUP_MEAN_STDEV_COLUMNS_ORDER +) ## Assert final column order includes all columns from original table if (!setequal(FINAL_COLUMN_ORDER, colnames(df_interim))) { @@ -1145,6 +1163,14 @@ get_versions <- function() { glue::glue(" homepage: https://www.r-project.org/"), glue::glue(" workflow task: PROCESS_AFFYMETRIX") ), sep = "\n") + # Add Bioconductor explicitly as a transitive dependency not captured by sessionInfo() + versions_buffer <- glue::glue_collapse(c( + versions_buffer, + glue::glue("- name: Bioconductor"), + glue::glue(" version: {packageVersion('BiocVersion')}"), + glue::glue(" homepage: https://bioconductor.org"), + glue::glue(" workflow task: PROCESS_AFFYMETRIX") + ), sep = "\n") # Get 'other attached packages' for (software in session_info[["otherPkgs"]]) { versions_buffer <- glue::glue_collapse(c( @@ -1168,23 +1194,15 @@ get_versions <- function() { return(versions_buffer) } -## Note Libraries that were NOT used during processing versions_buffer <- get_versions() -if (organism %in% c("athaliana")) { - versions_buffer <- glue::glue_collapse(c( - versions_buffer, - glue::glue("- name: biomaRt"), - glue::glue(" version: (Not used for plant datasets)"), - glue::glue(" homepage: https://bioconductor.org/packages/3.14/bioc/html/biomaRt.html"), - glue::glue(" workflow task: PROCESS_AFFYMETRIX") - ), sep = "\n") -} else if (organism %in% c('ecoli', 'paeruginosa')) { +## Note Libraries that were NOT used during processing +if (!grepl("purrr", versions_buffer)) { versions_buffer <- glue::glue_collapse(c( versions_buffer, - glue::glue("- name: biomaRt"), - glue::glue(" version: (Not used for bacteria datasets)"), - glue::glue(" homepage: https://bioconductor.org/packages/3.14/bioc/html/biomaRt.html"), + glue::glue("- name: purrr"), + glue::glue(" version: (Not used for this dataset)"), + glue::glue(" homepage: https://purrr.tidyverse.org/"), glue::glue(" workflow task: PROCESS_AFFYMETRIX") ), sep = "\n") } diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/GENERATE_SOFTWARE_TABLE/resources/usr/bin/SoftwareYamlToMarkdownTable.py b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/SoftwareYamlToMarkdownTable.py similarity index 69% rename from Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/GENERATE_SOFTWARE_TABLE/resources/usr/bin/SoftwareYamlToMarkdownTable.py rename to Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/SoftwareYamlToMarkdownTable.py index 76aa80c18..d1a4d7809 100755 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/GENERATE_SOFTWARE_TABLE/resources/usr/bin/SoftwareYamlToMarkdownTable.py +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/SoftwareYamlToMarkdownTable.py @@ -8,15 +8,15 @@ AFFYMETRIX_SOFTWARE_DPPD = [ "R", + "Bioconductor", "DT", "dplyr", "tibble", "stringr", - "R.utils", + "purrr", "oligo", "limma", "glue", - "biomaRt", "matrixStats", "statmod", "dp_tools", @@ -37,10 +37,12 @@ ## Used when the R library metadata doesn't encode any URLS HOMEPAGE_MAP = { "statmod":"https://cran.r-project.org/web/packages/statmod/index.html", - "biomaRt":"https://bioconductor.org/packages/3.14/bioc/html/biomaRt.html", # UPDATE ON biomaRt version update - "oligo":"https://www.bioconductor.org/packages/3.14/bioc/html/oligo.html", # UPDATE ON biomaRt version update + "oligo":"https://www.bioconductor.org/packages/3.22/bioc/html/oligo.html", # UPDATE ON bioconductor version update } +## Used when certain packages are conditionally used, and therefore dropped from the software table +NOT_USED_SENTINEL = "(Not used for this dataset)" + @click.command() @click.argument("input_yaml", type=click.Path(exists=True)) @click.argument("filename") @@ -53,13 +55,19 @@ def yamlToMarkdown(input_yaml: Path, filename: str, skip_de: bool): data.extend(ASSUMED_SOFTWARE) df = pd.DataFrame(data) - # If data files are not compressed, won't use R.utils to unzip them during processing - if not filename.endswith('.gz'): - AFFYMETRIX_SOFTWARE_DPPD.remove('r.utils') - if skip_de: AFFYMETRIX_SOFTWARE_DPPD.remove('limma') AFFYMETRIX_SOFTWARE_DPPD.remove('statmod') + AFFYMETRIX_SOFTWARE_DPPD.remove('matrixstats') + + # Drop software explicitly marked as unused for this dataset (e.g. purrr, only invoked on the 3prime-IVT custom-annotation branch) + # and remove them from the software table too, so the completeness assert below doesn't demand software that legitimately never ran + unused_mask = df["version"].astype(str) == NOT_USED_SENTINEL + unused_software = set(df.loc[unused_mask, "name"].str.lower()) + for name in unused_software: + if name in AFFYMETRIX_SOFTWARE_DPPD: + AFFYMETRIX_SOFTWARE_DPPD.remove(name) + df = df.loc[~unused_mask] # Filter to direct software used (i.e. exclude dependencies of the software) df = df.loc[df["name"].str.lower().isin(AFFYMETRIX_SOFTWARE_DPPD)] diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix/checks.py b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix/checks.py index fef5126e6..aca43c6ba 100644 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix/checks.py +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix/checks.py @@ -319,8 +319,8 @@ def utils_common_constraints_on_dataframe( col_constraints = col_constraints.copy() # limit to only columns of interest - query_df = df[col_set] - for (colname, colseries) in query_df.iteritems(): + query_df = df[list(col_set)] + for (colname, colseries) in query_df.items(): # check non null constraint if col_constraints.pop("nonNull", False) and nonNull(colseries) == False: issues["Failed non null constraint"].append(colname) @@ -398,7 +398,7 @@ def check_dge_table_sample_columns_constraints( ) -> FlagEntry: MINIMUM_COUNT = 0 # data specific preprocess - df_dge = pd.read_csv(dge_table)[samples] + df_dge = pd.read_csv(dge_table)[list(samples)] schema = pa.DataFrameSchema({ sample: pa.Column(float) diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix/config.yaml b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix/config.yaml index 071eff4f9..e718becda 100644 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix/config.yaml +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix/config.yaml @@ -1,6 +1,6 @@ # TOP LEVEL NAME: "microarray" -VERSION: "0" +VERSION: "1" # anchors for reuse _anchors: @@ -75,7 +75,10 @@ Staging: Sample name is used as a unique sample identifier during processing Example: Atha_Col-0_Root_WT_Ctrl_45min_Rep1_GSM502538 - - ISA Field Name: Label + - ISA Field Name: + - Label + - Parameter Value[label] + - Parameter Value[Label] ISA Table Source: Sample Runsheet Column Name: Label Processing Usage: >- @@ -176,7 +179,7 @@ data assets: runsheet: processed location: - "Metadata" - - "{dataset}_microarray_v0_runsheet.csv" + - "{dataset}_microarray_v1_runsheet.csv" tags: - raw @@ -263,16 +266,6 @@ data assets: resource categories: *DGEAnalysisData - viz PCA table: - processed location: - - *DGEDataDir - - "visualization_PCA_table_GLmicroarray.csv" - - tags: - - processed - - resource categories: *neverPublished - data asset sets: # These assets are not generated in the workflow, but are generated after the workflow PUTATIVE: [] diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix/protocol.py b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix/protocol.py index e17c9a1d2..d55f4f54f 100644 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix/protocol.py +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix/protocol.py @@ -365,33 +365,6 @@ def validate( """) ) - with vp.component_start( - name="Viz Tables", - description="Extended from the dge tables", - ): - with vp.payload( - payloads=[ - { - "samples": lambda: set(dataset.samples), - "pca_table": lambda: dataset.data_assets[ - "viz PCA table" - ].path, - } - ] - ): - vp.add( - bulkRNASeq.checks.check_viz_pca_table_index_and_columns_exist, - full_description=textwrap.dedent(f""" - - Check: Ensure all samples (row-indices) present and columns PC1, PC2 and PC3 are present - - Reason: - - PCA table should include all samples and PC1, PC2, PC3 (for 3D PCA viz) - - Potential Source of Problems: - - Bug in processing script - - Flag Condition: - - Green: All samples and all columns present - - Halt: At least one sample or column is missing - """) - ) with vp.component_start( name="Processing Report", description="", diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix_skipDE/checks.py b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix_skipDE/checks.py index fef5126e6..aca43c6ba 100644 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix_skipDE/checks.py +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix_skipDE/checks.py @@ -319,8 +319,8 @@ def utils_common_constraints_on_dataframe( col_constraints = col_constraints.copy() # limit to only columns of interest - query_df = df[col_set] - for (colname, colseries) in query_df.iteritems(): + query_df = df[list(col_set)] + for (colname, colseries) in query_df.items(): # check non null constraint if col_constraints.pop("nonNull", False) and nonNull(colseries) == False: issues["Failed non null constraint"].append(colname) @@ -398,7 +398,7 @@ def check_dge_table_sample_columns_constraints( ) -> FlagEntry: MINIMUM_COUNT = 0 # data specific preprocess - df_dge = pd.read_csv(dge_table)[samples] + df_dge = pd.read_csv(dge_table)[list(samples)] schema = pa.DataFrameSchema({ sample: pa.Column(float) diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix_skipDE/config.yaml b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix_skipDE/config.yaml index e30bf80a8..230edadbd 100644 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix_skipDE/config.yaml +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/dp_tools__affymetrix_skipDE/config.yaml @@ -1,6 +1,6 @@ # TOP LEVEL NAME: "microarray" -VERSION: "0" +VERSION: "1" # anchors for reuse _anchors: @@ -75,7 +75,10 @@ Staging: Sample name is used as a unique sample identifier during processing Example: Atha_Col-0_Root_WT_Ctrl_45min_Rep1_GSM502538 - - ISA Field Name: Label + - ISA Field Name: + - Label + - Parameter Value[label] + - Parameter Value[Label] ISA Table Source: Sample Runsheet Column Name: Label Processing Usage: >- @@ -176,7 +179,7 @@ data assets: runsheet: processed location: - "Metadata" - - "{dataset}_microarray_v0_runsheet.csv" + - "{dataset}_microarray_v1_runsheet.csv" tags: - raw diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/fetch_isa.py b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/fetch_isa.py new file mode 100755 index 000000000..436c63fbb --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/fetch_isa.py @@ -0,0 +1,72 @@ +#!/usr/bin/env python + +""" +Downloads the ISA archive for a given dataset using the provided OSD ID. +""" + +import argparse +import os +import sys +import requests + +def main(): + parser = argparse.ArgumentParser( + description="Download ISA archive for a given dataset using OSD ID." + ) + parser.add_argument("--osd", required=True, help="OSD ID (e.g., OSD-576)") + parser.add_argument("--outdir", required=True, help="Output directory to save the ISA archive") + args = parser.parse_args() + + outdir = args.outdir + + if not args.osd.startswith('OSD-'): + sys.exit(f"OSD accession ({args.osd}) was not provided in the correct format. It must start with 'OSD-'") + + # Build the JSON URL to get file information for a file with ISA in the name and a .zip extension + json_url = f"https://visualization.osdr.nasa.gov/biodata/api/v2/dataset/{args.osd}/files/" + + # Fetch the JSON data + response = requests.get(json_url) + if response.status_code != 200: + sys.exit(f"Error: Failed to retrieve file information. HTTP status code: {response.status_code}") + + # Parse the JSON data + data = response.json() + study_key = f"{args.osd}" + study_data = data.get(study_key) + if not study_data: + sys.exit(f"Error: Study with OSD ID '{args.osd}' not found in the response.") + + # Get the list of study files + files = study_data.get('files') + if not files: + sys.exit("Error: No files found in the response.") + + # Find the ISA archive file + try: + isa_file_name = next(filename for filename in files + if 'ISA' in filename and filename.endswith('.zip')) + except StopIteration: + sys.exit("Error: ISA archive not found in the file list.") + + # Construct the full download URL + download_url = files.get(isa_file_name).get('URL') + if not download_url: + sys.exit(f"Error: Download URL for ISA archive {isa_file_name} not found.") + + # Make the output directory and file path + os.makedirs(outdir, exist_ok=True) + output_file = os.path.join(outdir, isa_file_name) + + # Download the ISA archive + with requests.get(download_url, stream=True) as r: + if r.status_code == 200: + with open(output_file, 'wb') as f: + for chunk in r.iter_content(chunk_size=8192): + f.write(chunk) + print(f"ISA archive downloaded and saved to {output_file}") + else: + sys.exit(f"Error: Failed to download ISA archive. HTTP status code: {r.status_code}") + +if __name__ == "__main__": + main() \ No newline at end of file diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/generate_md5sum_files.py b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/generate_md5sum_files.py index 8faf0df39..035824008 100755 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/generate_md5sum_files.py +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/generate_md5sum_files.py @@ -1,63 +1,101 @@ -#! /usr/bin/env python -import argparse -from pathlib import Path +#!/usr/bin/env python3 -from dp_tools.core.loaders import load_data -from dp_tools.core.post_processing import generate_md5sum_table -from dp_tools.plugin_api import load_plugin +import os +import sys +import hashlib +import argparse +def calculate_md5(filepath): + """Calculate MD5 hash for a file.""" + md5_hash = hashlib.md5() + + # Follow symlinks to get the actual file + actual_path = os.path.realpath(filepath) if os.path.islink(filepath) else filepath + + try: + with open(actual_path, "rb") as f: + # Read in chunks in case of large files + for chunk in iter(lambda: f.read(4096), b""): + md5_hash.update(chunk) + return md5_hash.hexdigest() + except Exception as e: + sys.stderr.write(f"Error calculating MD5 for {filepath}: {str(e)}\n") + return "ERROR" -############################################################## -# Utility Functions To Handle Logging, Config and CLI Arguments -############################################################## -def _parse_args(): - """Parse command line args.""" - parser = argparse.ArgumentParser() +def should_include(filepath): + """Check if file should be included in MD5 calculation.""" + # Skip files in GeneLab except for HTML report, software_versions, and purged processing_info.txt + allowed_files = [ + "NF_MAAffymetrix_v" + args.workflow_version + "_GLmicroarray.html", + "software_versions_GLmicroarray.md", + "nextflow_processing_info_GLmicroarray.txt" + ] + if "/GeneLab/" in filepath and not any(filepath.endswith(f) for f in allowed_files): + return False + + # Skip ISA.zip + if filepath.endswith("ISA.zip"): + return False - parser.add_argument("--root-path", required=True, help="Root data path") + # Skip VV logs + if "/VV_Logs/" in filepath: + return False + + return True - parser.add_argument("--runsheet-path", required=True, help="Runsheet path") +def main(): + parser = argparse.ArgumentParser(description='Generate MD5 sum files for GeneLab data.') + parser.add_argument('--outdir', required=True, help='Output directory containing files to process') + parser.add_argument('--workflow_version', default='', help='Version of the NF_MAAffymetrix workflow manifest') + + global args + args = parser.parse_args() + + # Make sure outdir is absolute path + outdir = os.path.abspath(args.outdir) + + # Create output files and initialize them without headers + processed_md5_file = f"processed_md5sum_GLmicroarray.tsv" + with open(processed_md5_file, 'w') as f: + f.write("File Name\tmd5sum\n") + + + # Track processed files for reporting + processed_count = 0 + + # Walk through all files recursively + print(f"Scanning directory: {outdir}") + for root, _, files in os.walk(outdir): + for filename in files: + filepath = os.path.join(root, filename) + + # Skip files that shouldn't be included + if not should_include(filepath): + continue + + # Get just the filename (basename) + basename = os.path.basename(filepath) + + md5sum = calculate_md5(filepath) + with open(processed_md5_file, 'a') as f: + f.write(f"{basename}\t{md5sum}\n") + processed_count += 1 + + print(f"Added {processed_count} files to {processed_md5_file}") - parser.add_argument("--plug-in-dir", required=True, help="Plugin path") + def dedup_file(filename): + seen = set() + lines = [] + with open(filename, 'r') as f: + for line in f: + key = line.split('\t', 1)[0] # dedup by basename + if key not in seen: + seen.add(key) + lines.append(line) + with open(filename, 'w') as f: + f.writelines(lines) - args = parser.parse_args() - return args - - -def main(root_dir: Path, runsheet_path: Path, plug_in_dir: Path): - plugin = load_plugin(Path(plug_in_dir)) - - ds = load_data( - config=plugin.config, - root_path=(root_dir), - runsheet_path=runsheet_path, - ) - - df = generate_md5sum_table( - ds.dataset, - config=plugin.config, - include_tags=True, - ) - - unique_tags = set(df["tags"].sum()) - for tag in unique_tags: - df_subset = df.loc[df["tags"].apply(lambda l: tag in l)].drop( - "tags", axis="columns" - ) - df_subset.to_csv(f"{tag}_md5sum_GLmicroarray.tsv", sep="\t", index=False) - - # Log missing files - print(df.columns) - missing_files = df.loc[df['md5sum'] == "USER MUST ADD MANUALLY!"]["filename"].to_list() - if missing_files: - with open("Missing_md5sum_files_GLmicroarray.txt", "w") as f: - for missing in missing_files: - f.write(missing+"\n") + dedup_file(processed_md5_file) if __name__ == "__main__": - import logging - - logging.basicConfig(level=logging.DEBUG) - log = logging.getLogger(__name__) - args = _parse_args() - main(Path(args.root_path), runsheet_path=Path(args.runsheet_path), plug_in_dir=args.plug_in_dir) \ No newline at end of file + main() \ No newline at end of file diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/generate_protocol.py b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/generate_protocol.py new file mode 100755 index 000000000..173a161b8 --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/generate_protocol.py @@ -0,0 +1,326 @@ +#!/usr/bin/env python +""" +This script generates a protocol text file for GeneLab Affymetrix Microarray data processing. +It reads software versions from a YAML file and incorporates other parameters. +""" + +import argparse +import os +import sys +import re +from datetime import datetime +from pathlib import Path +import pandas as pd +import yaml + +def parse_args(): + """Sets up and parses the input parameters using argparse + + Returns: + NameSpace: object holding input parameters + """ + parser = argparse.ArgumentParser(description='Generate protocol file for GeneLab Affymetrix Microarray pipeline') + parser.add_argument('--outdir', required=True, type=Path, + help='Output directory for the protocol file') + parser.add_argument('--software_table', required=True, type=Path, + help='Path to YAML file containing software versions') + parser.add_argument('--assay_suffix', default='', + help='Suffix for the GeneLab assay type') + parser.add_argument('--workflow_version', default='unknown', + help='Version of the NF_MAAffymetrix workflow manifest') + parser.add_argument('--organism', required=False, + help='Organism name in the format "homo_sapiens"') + parser.add_argument('--reference_source', required=False, + help='Source of the reference annotation') + parser.add_argument('--reference_version', required=False, + help='Version of the reference annotation') + parser.add_argument('--biomart_attribute', required=False, + help='Attribute for biomart query (if applicable)') + parser.add_argument('--bioconductor_annotations', required=False, type=str, + help='bioconductor annotation database name') + parser.add_argument('--annotations_db_info', required=False, type=Path, + help='Annotation DB info file from GeneLab Reference Annotation database used') + parser.add_argument('--custom_annot_design', required=False, type=Path, + help='Path to the custom probe annotation design info file') + parser.add_argument('--skip-DGE', type=bool, required=False, + help="Was DGE performed.") + return parser.parse_args() + + +def read_annotation_versions(annot_info_file: Path, bioc_annot: str) -> dict[str, str]: + """Reads the version information from the annotation DB info file into a dictionary mapping software name to version + + Args: + annot_info_file (Path): the annotation info file produced by the GeneLab Reference Annotation pipeline + + Returns: + dict: a dictionary mapping software name to version + """ + versions = dict() + try: + with open(annot_info_file, 'r') as f: + bioc_annot_found = False + software_name = '' + version_found = False + doc_name_found = False + for line in f: + if bioc_annot_found: + versions[bioc_annot] = line.strip() + bioc_annot_found = False + if version_found: + versions[software_name] = line.strip() + software_name = '' + version_found = False + if doc_name_found: + versions['annot_doc'] = line.strip() + doc_name_found = False + if line.startswith('Used'): + if re.search(bioc_annot, line): + bioc_annot_found = True + else: + version_found = True + software_name = line.split(' ')[1] + if line.startswith('Based on:'): + doc_name_found = True + except Exception as e: + sys.stderr.write(f"Error reading annotation info file: {e}\n") + sys.exit(1) + + return versions + + +def read_software_versions(yaml_file: Path) -> dict: + """Reads the software versions YAML file into a dictionary mapping software name to version + + Args: + yaml_file (Path): a YAML formatted file containing software version info produced by a Nextflow workflow + + Returns: + dict: a dictionary mapping software name to version + """ + try: + with open(yaml_file, 'r') as f: + return pd.DataFrame(yaml.safe_load(f)).set_index('name').to_dict()['version'] + except Exception as e: + sys.stderr.write(f"Error reading software versions file: {e}\n") + sys.exit(1) + + +def create_header(assay_suffix, workflow_version): + """Create protocol header + + Args: + assay_suffix (str): GeneLab assay suffix + workflow_version (str): Workflow version + + Returns: + str: Protocol header section + """ + # Get current date + current_date = datetime.now().strftime("%Y-%m-%d") + + # Create header + header = f"# GeneLab Microarray Pipeline Protocol{assay_suffix}\n" + header += f"# Date: {current_date}\n\n" + + # Add appropriate protocol reference based on mode + header += "Data were processed as described in GL-DPPD-7114-A " + header += "(https://github.com/nasa/GeneLab_Data_Processing/blob/master/Microarray/Affymetrix/Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114-A.md), " + + header += f"using NF_MAAffymetrix version {workflow_version} " + header += f"(https://github.com/nasa/GeneLab_Data_Processing/tree/NF_MAAffymetrix_{workflow_version}/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix)." + + return header + + +def create_annot_and_dge_section(organism, biomart_attribute, ensembl_version, + bioconductor_annotations, r_version, limma_version, annotation_versions, + custom_annots=pd.DataFrame(), skip_DGE=False): + """Generate the probe annotation and DGE protocol sections + + Args: + organism (str): full organism name lowercase with underscores instead of spaces (e.g., "homo_sapiens") + biomart_attribute (str): Attribute for biomart query (name of the Agilent array as present in biomart or the custom design info file, if applicable) + ensembl_version (str): Ensembl reference version + bioconductor_annotations (str): Bioconductor annotation package name + r_version (str): R software version + limma_version (str): limma software version + custom_annots (pd.DataFrame, optional): path to custom probe annotation design info file used during processing. Defaults to an empty path + skip_DGE (bool, optional): Whether or not DGE was skipped during processing. Defaults to False. + + Returns: + str: Protocol string for probe annotation and DGE + """ + + custom_annot_source_name = None + custom_annot_download_link = None + custom_annot_download_date = None + custom_annot_create_date = None + custom_annot_filename = None + + if not custom_annots.empty and biomart_attribute in custom_annots.index.values: + custom_annot_source_name = custom_annots['annot_type'][biomart_attribute] + custom_annot_filename = custom_annots['annot_filename'][biomart_attribute] + if 'download_link' in custom_annots.columns: + custom_annot_download_link = custom_annots['download_link'][biomart_attribute] + custom_annot_download_date = custom_annots['download_date'][biomart_attribute] + if 'create_date' in custom_annots.columns: + custom_annot_create_date = custom_annots['create_date'][biomart_attribute] + + # Define versions for annotation package generation + annot_doc = annotation_versions['annot_doc'] if 'annot_doc' in annotation_versions else 'GL-DPPD-7110-A' + stringdb_version = annotation_versions['STRINGdb'] if 'STRINGdb' in annotation_versions else "2.16.4" + pantherdb_version = annotation_versions['PANTHER.db'] if 'PANTHER.db' in annotation_versions else "1.0.12" + + # Define organism to annotation package mapping using scientific names + organism_annotation_package = annotation_versions[bioconductor_annotations] if bioconductor_annotations in annotation_versions else "3.19.1" + + # Check if DGE was performed + de_step = "" + if not skip_DGE: + de_step = f"Differential expression analysis was performed in R (version {r_version}) using limma (version {limma_version}); " + de_step += "all groups were compared pairwise for each probeset to generate a moderated t-statistic and associated p- and adjusted p-value." + + organism_list=("homo_sapiens", "mus_musculus", "rattus_norvegicus", "drosophila_melanogaster", "caenorhabditis_elegans", "danio_rerio", "saccharomyces_cerevisiae") + gene_mapping_step = "" + + # Case 1: a custom annotation source was used for this array design + # no Ensembl FTP lookup happens in this case, which never calls + # the Ensembl FTP helpers when annot_type is 'custom' or '3prime-IVT'. + if biomart_attribute in custom_annots.index.values and custom_annot_source_name is not None: + if '3prime-IVT' in custom_annot_source_name: + gene_mapping_step = f"Gene annotations " + annot_source = "3'-IVT" + else: + gene_mapping_step = "Annotations " + annot_source = custom_annot_source_name.replace('_', ' ').title() + gene_mapping_step += f"were retrieved for each probeset from {custom_annot_filename}, source: {annot_source}" + if custom_annot_download_link is not None: + created_text = "" + if custom_annot_create_date is not None: + created_text = f"created {custom_annot_create_date}, " + gene_mapping_step += f" ({custom_annot_download_link}, {created_text}accessed {custom_annot_download_date})." + else: + gene_mapping_step += "." + + # Case 2: Ensembl FTP mart dump was used (no custom annotation) + else: + if organism == "arabidopsis_thaliana": + database_name = "Plants Ensembl database" + database_url = "plants.ensembl.org" + elif organism in organism_list: + database_name = "Ensembl database" + database_url = "ensembl.org" + else: # what should we do if the organism is not in the list?? + database_name = "TBD" + database_url = "TBD" + gene_mapping_step = f"Ensembl gene ID mappings were retrieved for each probeset using the {database_name} ftp server ({database_url}, release {ensembl_version})." + + # Gene annotations (STRINGdb/PANTHER/bioconductor merge) only runs when the + # Ensembl FTP path succeeded (use_custom_annot == False in the QMD); every + # custom-annotation branch (3prime-IVT, custom, NO_CUSTOM_ANNOT) skips this. + annot_step = "" + if custom_annot_source_name is None: + annot_step = f"Gene annotations were assigned using the custom annotation tables generated in-house as detailed in {annot_doc} " + annot_step += f"(https://github.com/nasa/GeneLab_Data_Processing/blob/master/GeneLab_Reference_Annotations/Pipeline_GL-DPPD-7110_Versions/{annot_doc}/{annot_doc}.md), " + annot_step += f"with STRINGdb (version {stringdb_version}), PANTHER.db (version {pantherdb_version}), and {bioconductor_annotations} (version {organism_annotation_package})." + + return " ".join((gene_mapping_step, de_step, annot_step)) + + +def generate_protocol_content(args:argparse.Namespace, software_versions:dict, annotation_versions:dict, assay_suffix:str = "") -> str: + """Generates a protocol string based on the input parameters and software versions + + Args: + args (argparse): script input parameters + software_versions (dict): a dictionary mapping software name to version + assay_suffix (str): the Genelab assay suffix + + Returns: + str: protocol text + """ + + header = create_header(assay_suffix, args.workflow_version) + + # Start building the description as a single paragraph + # Add processing description with software versions + oligo_version = software_versions.get('oligo', 'unknown') + + custom_annots = pd.DataFrame() + if args.custom_annot_design.exists() and args.custom_annot_design.is_file(): + custom_annots = pd.read_csv(open(args.custom_annot_design, "r")).set_index('array_design') + + description = f"In short, a runsheet containing raw data file location and processing metadata from the study's *ISA.zip file was generated using dp_tools (version {software_versions.get('dp_tools', 'unknown')}). " + description += f"The raw array data files were loaded into R (version {software_versions.get('R', 'unknown')}) using oligo (version {oligo_version}). " + description += f"Raw data quality assurance density, pseudo image, MA, and boxplots were generated using oligo (version {oligo_version}). " + description += f"The raw intensity data was background corrected and normalized across arrays via the oligo (version {oligo_version}) quantile method. " + description += f"Normalized data quality assurance density, pseudo image, MA plots, and boxplots were generated using oligo (version {oligo_version}). " + description += f"Normalized probe level data was summarized to the probeset level using the oligo (version {oligo_version}) RMA method." + + R_version = software_versions.get('R', 'unknown') + limma_version = software_versions.get('limma', 'unknown') + + annot_and_dge_section = create_annot_and_dge_section(args.organism, args.biomart_attribute, + args.reference_version, args.bioconductor_annotations, + R_version, limma_version, + annotation_versions, custom_annots, args.skip_DGE) + + # Create configuration section + config = "\n\n\n## Configuration\n\n" + + # Add reference information if provided + if hasattr(args, 'organism') and args.organism: + config += f"- Organism: {args.organism}\n" + if hasattr(args, 'reference_source') and args.reference_source: + config += f"- Reference source: {args.reference_source}\n" + if hasattr(args, 'reference_version') and args.reference_version: + config += f"- Reference version: {args.reference_version}\n" + + # Add custom reference information if provided + if hasattr(args, 'biomart_attribute') and args.biomart_attribute in custom_annots.index.values: + config += f"- Array design: {args.biomart_attribute}\n" + config += f"- Custom annotation source: {custom_annots['annot_type'][args.biomart_attribute]}\n" + config += f"- Custom annotation file: {custom_annots['annot_filename'][args.biomart_attribute]}\n" + if custom_annots['download_link'][args.biomart_attribute] is not None: + config += f"- Custom annotation download link: {custom_annots['download_link'][args.biomart_attribute]}\n" + config += f"- Custom annotation download date: {custom_annots['download_date'][args.biomart_attribute]}\n" + + # Create software versions section + sw_section = "\n\n## All Software Versions\n\n" + for software, version in software_versions.items(): + sw_section += f"- {software}: {version}\n" + + # Combine all sections + return header + " " + description + " " + annot_and_dge_section + config + sw_section + + +def main(): + args = parse_args() + + # Read software versions from YAML file + software_versions = read_software_versions(args.software_table) + + annotation_versions = dict() + if args.annotations_db_info is not None and args.bioconductor_annotations is not None: + annotation_versions = read_annotation_versions(args.annotations_db_info, args.bioconductor_annotations) + + assay_suffix = args.assay_suffix + if not assay_suffix.startswith('_'): + assay_suffix = f"_{assay_suffix}" + + # Generate protocol content + protocol_content = generate_protocol_content(args, software_versions, annotation_versions, assay_suffix) + + + # Write to output file + output_file = os.path.join(args.outdir, f"protocol{assay_suffix}.txt") + try: + with open(output_file, 'w') as f: + f.write(protocol_content) + print(f"Protocol file generated successfully: {output_file}") + except Exception as e: + sys.stderr.write(f"Error writing protocol file: {e}\n") + sys.exit(1) + +if __name__ == "__main__": + main() diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/get_accessions.py b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/get_accessions.py new file mode 100755 index 000000000..5c486ac37 --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/get_accessions.py @@ -0,0 +1,80 @@ +#!/usr/bin/env python + +import requests +import argparse +import re +import sys +import json + +def get_osd_and_glds(accession, api_url): + # Fetch data from the API + try: + response = requests.get(api_url) + response.raise_for_status() + data = response.json() + except requests.exceptions.RequestException as e: + print(f"Error fetching data from API: {e}", file=sys.stderr) + sys.exit(1) + except json.JSONDecodeError: + print("Error decoding JSON response from API", file=sys.stderr) + sys.exit(1) + + osd_accession = None + glds_accessions = [] + + # Check if the accession is OSD or GLDS + if accession.startswith('OSD-'): + osd_accession = accession + + # Find this OSD in the data + # Search in wildcard endpoint results + for osd_id, osd_data in data.items(): + if osd_id == accession: + metadata = osd_data.get("metadata", {}) + identifiers = metadata.get("identifiers", []) + # Normalize to list to handle single and multiple entries the same way + if isinstance(identifiers, str): + identifiers = [identifiers] + elif not isinstance(identifiers, list): + identifiers = [] + + glds_accessions = [ + x for x in identifiers + if isinstance(x, str) and x.startswith("GLDS-") + ] + break + + elif accession.startswith('GLDS-'): + glds_accessions = [accession] + + # Find the OSD associated with this GLDS + for osd_id, osd_data in data.items(): + metadata = osd_data.get("metadata", {}) + identifiers = metadata.get("identifiers", "") + if accession in identifiers: + osd_accession = metadata.get("accession") + break + else: + print("Invalid accession format. Please use 'OSD-###' or 'GLDS-###'.", file=sys.stderr) + sys.exit(1) + + if not osd_accession or not glds_accessions: + print(f"No data found for {accession}", file=sys.stderr) + sys.exit(1) + + return osd_accession, glds_accessions + +def main(): + parser = argparse.ArgumentParser(description="Retrieve OSD and GLDS accessions.") + parser.add_argument('--accession', required=True, help="Accession in the format 'OSD-###' or 'GLDS-###'") + parser.add_argument('--api_url', required=True, help="OSDR API URL") + args = parser.parse_args() + + osd_accession, glds_accessions = get_osd_and_glds(args.accession, args.api_url) + + # Output the results in a way that Nextflow can capture + print(f"{osd_accession}") + print(f"{','.join(glds_accessions)}") + +if __name__ == "__main__": + main() \ No newline at end of file diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/DUMP_META/resources/usr/bin/reformat_meta.sh b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/reformat_meta.sh similarity index 100% rename from Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/DUMP_META/resources/usr/bin/reformat_meta.sh rename to Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/reformat_meta.sh diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/update_assay_table.py b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/update_assay_table.py new file mode 100755 index 000000000..8549b9eff --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/update_assay_table.py @@ -0,0 +1,509 @@ +#!/usr/bin/env python + +import sys +import argparse +import zipfile +import pandas as pd +import json +import re + + +def parse_args(): + parser = argparse.ArgumentParser( + prog='update_assay_table', + description='Update Microarray Affymetrix assay table from ISA.zip with processed data file information.') + required = parser.add_argument_group('Required arguments') + required.add_argument('--runsheet', required=True, + help='Runsheet') + required.add_argument('--glds_accession', required=True, + help='GLDS accession number (e.g. GLDS-123)') + required.add_argument('--isa_zip', action='store', default='', + help='Appropriate ISA file for the dataset') + return parser.parse_args() + + +tty_colors = { + 'green': '\033[0;32m%s\033[0m', + 'yellow': '\033[0;33m%s\033[0m', + 'red': '\033[0;31m%s\033[0m' +} + + +def color_text(text, color='green'): + """ + Colors text for output in terminal + + Args: + text (str): input text + color (str): a valid tty color ('red', 'yellow', or 'green') + + Returns: + str: colored text + + """ + if sys.stdout.isatty(): + return tty_colors[color] % text + else: + return text + + +def report_failure_and_exit(message, color="red"): + """ + Reports a failure and exits with status '1'. + + Args: + message (str): Error message to report. + color (str): Color in which to render the error message, default = 'red' + """ + print("") + print(color_text(f"Error: {message}", color)) + print("\nAssay table update failed.\n") + + sys.exit(1) + + +def report_warning(message, color="yellow"): + """ + Reports are warning message. + + Args: + message (str): Error message to report + color (str): Color in which to render the error message, default = 'yellow' + """ + print("") + print(color_text(f"Warning: {message}", color)) + + +def load_runsheet(runsheet_file): + """ + Load the runsheet as a pandas.DataFrame. + + Args: + runsheet_file (PathLike[str]): a file containing the assay runsheet used to generate the processed data + + Returns: + pandas.DataFrame: sample information from the runsheet + """ + try: + runsheet_df = pd.read_csv(runsheet_file) + print(f"Runsheet has {len(runsheet_df)} rows and {len(runsheet_df.columns)} columns") + return runsheet_df + except Exception as e: + report_warning(f"Cannot read runsheet, proceeding without it: {e}") + return None + + +def get_runsheet_sample_name_map(runsheet_df, assay_sample_names): + """ + Generates a mapping of sample names in the assay table to the samplenames in the runsheet + + Args: + runsheet_df (pandas.DataFrame): runsheet sample information + assay_sample_names (list): sample names from the assay table + + Returns: + dict: sample name mapping + """ + sample_name_map = {} + if 'Sample Name' in runsheet_df.columns: + # Check for 'Original Sample Name' column to map between assay table and runsheet + if 'Original Sample Name' in runsheet_df.columns: + for _, row in runsheet_df.iterrows(): + orig_name = row['Original Sample Name'] + rs_name = row['Sample Name'] + if orig_name in assay_sample_names: + sample_name_map[orig_name] = rs_name + return sample_name_map + + + +def get_assay_table_from_isa(isa_file): + """ + tries to find an assay table in an ISA zip file that matches the type expected for the provided assay + + Args: + isa_file (PathLike[str]): path to ISA zip file + + Returns: + pandas.DataFrame: assay table from extracted from ISA zip + """ + valid_measurement = "transcription profiling" + valid_technology = "DNA microarray" + valid_platform = "Affymetrix" + + zip_file = zipfile.ZipFile(isa_file) + isa_files = zip_file.namelist() + + # Parse investigation file to build STUDY ASSAYS table + study_assays_table = {} + study_assays_section = False + + investigation_file = next((f for f in isa_files if f.startswith('i_') and f.endswith('.txt')), None) + if not investigation_file: + report_failure_and_exit(f"Investigation file not found in ISA zip: {isa_file}") + + with zip_file.open(investigation_file, 'r') as f: + for line in f: + line = line.decode('utf-8').strip() + + # Track STUDY ASSAYS section + if line == 'STUDY ASSAYS': + study_assays_section = True + continue + elif study_assays_section and not line: + study_assays_section = False + continue + + # Extract data from section + if study_assays_section and line: + parts = line.split('\t') + if parts and parts[0]: + key = parts[0] + values = [v.strip() for v in parts[1:] if v.strip()] + study_assays_table[key] = values + # Check if we have all required keys + required_keys = ['Study Assay Measurement Type', 'Study Assay Technology Type', 'Study Assay Technology Platform', 'Study Assay File Name'] + if not all(key in study_assays_table for key in required_keys): + report_failure_and_exit("Missing required keys in STUDY ASSAYS section") + + # Get the values from the table + measurement_types = study_assays_table['Study Assay Measurement Type'] + technology_types = study_assays_table['Study Assay Technology Type'] + technology_platforms = study_assays_table['Study Assay Technology Platform'] + file_names = study_assays_table['Study Assay File Name'] + + # Ensure all lists have equal length + if not (len(measurement_types) == len(technology_types) == len(technology_platforms) == len(file_names)): + report_failure_and_exit("Measurement types, technology types, technology platforms, and file names have different lengths") + + # Find matching assay file + matched_file = "" + for i in range(len(measurement_types)): + if (measurement_types[i].lower() == valid_measurement.lower() and + technology_types[i].lower() == valid_technology.lower() and + technology_platforms[i].lower() == valid_platform.lower()): + matched_file = file_names[i] + break + + if not matched_file: + report_failure_and_exit(f"No assay file matched for {valid_measurement} assay. " + f"Measurement types: {measurement_types}, Technology types: {technology_types}, Technology platforms: {technology_platforms}") + elif matched_file not in isa_files: + # Load the matched assay file + report_failure_and_exit(f"Matched assay file doesn't exist in ISA zip: {matched_file}") + else: + return pd.read_csv(zip_file.open(matched_file), sep='\t'), matched_file + + return pd.DataFrame(), matched_file + + +def add_parameter_column(df, column_name, value, glds_prefix=None): + """ + Add a parameter column to the dataframe if it doesn't exist already. Use the same value for all rows in the table. + + Args: + df (pandas.DataFrame): current assay table + glds_prefix (str): a prefix to add to the start of each value (only for values that are filenames) + column_name (str): parameter column name to add (e.g., "Parameter Value[Entry]") + value (str): value to set for all rows in the table + + Returns: + pandas.DataFrame: update assay table + """ + # Apply prefix to value if provided + if glds_prefix and isinstance(value, str): + # Check if value already has the prefix + if not value.startswith(glds_prefix): + prefixed_value = f"{glds_prefix}{value}" + else: + prefixed_value = value + else: + prefixed_value = value + + if column_name not in df.columns: + print(f"Adding new column: {column_name}") + df[column_name] = prefixed_value + else: + print(f"Column {column_name} already exists, updating values") + df[column_name] = prefixed_value + + return df + + +def add_protocol_ref_column(df): + """Add Protocol REF GeneLab microarray data processing protocol column""" + value = "GeneLab microarray data processing protocol" + + # Insert after "Array Data File" or "Derived Array Data File" if present, otherwise at the end + insert_position = None + + # Pass 1: prefer the first "Derived Array Data File" column + for i, col in enumerate(df.columns): + if "Derived Array Data File" in col: + column = "Derived Array Data File" + insert_position = i + 1 + break + + # Pass 2: fall back to the first "Array Data File" column + if insert_position is None: + for i, col in enumerate(df.columns): + if "Array Data File" in col: + column = "Array Data File" + insert_position = i + 1 + break + + # Pass 3: fall back to end of dataframe + if insert_position is None: + column = "end of dataframe" + insert_position = len(df.columns) + + # Remove any existing Protocol REF columns with the exact data processing protocol value + drop_cols = [] + drop_indices = [] + for i, col in enumerate(df.columns): + if "protocol ref" in col.lower(): + col_data = df.iloc[:, i].astype(str).str.strip() + if (col_data.str.lower() == value.lower()).any(): + drop_cols.append(col) + drop_indices.append(i) + + if drop_cols: + print(f"Removing existing Protocol REF columns with data processing protocol: {', '.join(drop_cols)}") + df = df.drop(columns=drop_cols) + # Adjust insert position if needed + removed_before = sum(1 for idx in drop_indices if idx < insert_position) + insert_position -= removed_before + if insert_position < 0: + insert_position = 0 + + # Insert with temp name, rename "Protocol REF" + temp_name = "Protocol REF_DP" + while temp_name in df.columns: + temp_name = f"{temp_name}_temp" + df.insert(insert_position, temp_name, value) + cols = list(df.columns) + cols[insert_position] = "Protocol REF" + df.columns = cols + + print(f"Added: Protocol REF (value: {value}) after {column}") + return df + + +def add_raw_intensities_table_column(df, glds_prefix): + """ + Add or update the raw intensities table column to the dataframe. + + Args: + df (pandas.DataFrame): current assay table + glds_prefix (str): a prefix to add to the start of each filename + + Returns: + pandas.DataFrame: updated assay table + """ + column_name = "Parameter Value[Raw Intensities Table]" + # Create the raw intensities filename - same for all samples + raw_intensities = (f"{glds_prefix}_array_raw_intensities_probes_GLmicroarray.csv") + + # Look for an existing column matching this name, ignoring case, + # so we update/rename it instead of creating a duplicate column + existing_col = next((col for col in df.columns if col.lower() == column_name.lower()), None) + + # If a differently-cased version exists, rename it to the canonical name + if existing_col and existing_col != column_name: + print(f"Renaming column '{existing_col}' to '{column_name}'") + df = df.rename(columns={existing_col: column_name}) + + # Set the value for all rows (creates the column if it didn't exist) + print(f"{'Updating' if existing_col else 'Adding'} column: {column_name}") + df[column_name] = raw_intensities + + return df + + +def add_normalized_expression_table_column(df, glds_prefix): + """ + Add or update the normalized expression table column to the dataframe. + + Args: + df (pandas.DataFrame): current assay table + glds_prefix (str): a prefix to add to the start of each filename + + Returns: + pandas.DataFrame: updated assay table + """ + column_name = "Parameter Value[Normalized Expression Table]" + # Create the normalized expression filenames - same for all samples + normalized_files = [ + f"{glds_prefix}_array_normalized_expression_probset_GLmicroarray.csv", + f"{glds_prefix}_array_normalized_intensities_probe_GLmicroarray.csv" + ] + + combined_files = ','.join(normalized_files) + + # Look for an existing column matching this name, ignoring case, + # so we update/rename it instead of creating a duplicate column + existing_col = next((col for col in df.columns if col.lower() == column_name.lower()), None) + + # If a differently-cased version exists, rename it to the canonical name + if existing_col and existing_col != column_name: + print(f"Renaming column '{existing_col}' to '{column_name}'") + df = df.rename(columns={existing_col: column_name}) + + # Set the value for all rows (creates the column if it didn't exist) + print(f"{'Updating' if existing_col else 'Adding'} column: {column_name}") + df[column_name] = combined_files + + return df + + +def add_differential_expression_analysis_data_column(df, glds_prefix): + """ + Add or update the Differential Expression Analysis Data column to the dataframe. + + Args: + df (pandas.DataFrame): current assay table + glds_prefix (str): a prefix to add to the start of each filename + + Returns: + pandas.DataFrame: updated assay table + """ + column_name = "Parameter Value[Differential Expression Analysis Data]" + # Create the differential expression filenames - same for all samples + de_files = [ + f"{glds_prefix}_array_SampleTable_GLmicroarray.csv", + f"{glds_prefix}_array_contrasts_GLmicroarray.csv", + f"{glds_prefix}_array_differential_expression_GLmicroarray.csv" + ] + + # Join the files with commas + combined_files = ','.join(de_files) + + # Look for an existing column matching this name, ignoring case, + # so we update/rename it instead of creating a duplicate column + existing_col = next((col for col in df.columns if col.lower() == column_name.lower()), None) + + # If a differently-cased version exists, rename it to the canonical name + if existing_col and existing_col != column_name: + print(f"Renaming column '{existing_col}' to '{column_name}'") + df = df.rename(columns={existing_col: column_name}) + + # Set the value for all rows (creates the column if it didn't exist) + print(f"{'Updating' if existing_col else 'Adding'} column: {column_name}") + df[column_name] = combined_files + + return df + + +def clean_comma_space(df): + """Remove spaces after commas in all string columns of the dataframe.""" + for col in df.columns: + try: + # If there are duplicate columns, df[col] is a DataFrame, not Series + if hasattr(df[col], 'dtype') and df[col].dtype == 'object': + df[col] = df[col].str.replace(", ", ",", regex=False) + except Exception: + # If df[col] is a DataFrame (duplicate columns), apply to each + for c in range(df.columns.get_loc(col), len(df.columns)): + if df.columns[c] == col: + s = df.iloc[:, c] + if s.dtype == 'object': + df.iloc[:, c] = s.str.replace(", ", ",", regex=False) + print("Removed spaces after commas in all string columns") + return df + + +def clean_column_names(df): + """Clean column names by removing any .# suffixes pandas adds to duplicates. + + Args: + df: The DataFrame to clean column names + + Returns: + The DataFrame with cleaned column names + """ + # Create a mapping of old_name -> new_name (without .# suffix) + name_mapping = {} + for col in df.columns: + # Use regex to match column names with .digits suffix + if re.search(r'\.\d+$', col): + # Remove the .# suffix + base_name = re.sub(r'\.\d+$', '', col) + name_mapping[col] = base_name + + # Rename columns using the mapping if any found + if name_mapping: + print(f"Cleaning {len(name_mapping)} column names by removing .# suffixes:") + for old_name, new_name in name_mapping.items(): + print(f" - {old_name} -> {new_name}") + df = df.rename(columns=name_mapping) + + return df + + +def main(): + args = parse_args() + + # Find and parse the runsheet + runsheet_df = load_runsheet(args.runsheet) + + # Find Microarray sequencing assay file and get its contents - use assay table if directly provided, else extract it from ISA.zip + print(f"Extracting assay table from {args.isa_zip}") + assay_df, assay_filename = get_assay_table_from_isa(args.isa_zip) + print(f"Original assay table has {len(assay_df)} rows and {len(assay_df.columns)} columns") + + + # Create a mapping from assay table sample names to runsheet sample names, exit if no sample column found + sample_col = next((col for col in assay_df.columns if 'Sample Name' in col), None) + if sample_col is None: + report_failure_and_exit(f"Could not find 'Sample Name' column in assay table '{assay_filename}'") + + assay_sample_names = assay_df[sample_col].tolist() + if runsheet_df is not None and 'Sample Name' in runsheet_df.columns: + sample_name_map = get_runsheet_sample_name_map(runsheet_df, assay_sample_names) + else: + sample_name_map = {} + + # Process and save assay file + try: + + + assay_df = add_protocol_ref_column(assay_df) + + # Raw Intensities Table column + assay_df = add_raw_intensities_table_column(assay_df, args.glds_accession) + + # Normalized Expression Table column + assay_df = add_normalized_expression_table_column(assay_df, args.glds_accession) + + # Differential Expression Analysis Data column + assay_df = add_differential_expression_analysis_data_column(assay_df, args.glds_accession) + + # Clean comma-space in all string columns + assay_df = clean_comma_space(assay_df) + + # Clean column names by removing any .# suffixes pandas adds + assay_df = clean_column_names(assay_df) + + # Use the filename we found in extract_and_find_assay + orig_filename = assay_filename + + # Create both original and modified output files + # Original file (preserving the original name) + assay_df.to_csv(orig_filename, sep='\t', index=False) + print(f"Original assay table saved as: {orig_filename}") + + # Modified file with GLDS prefix + if not orig_filename.startswith(args.glds_accession): + mod_filename = f"{args.glds_accession}{orig_filename}" + else: + mod_filename = orig_filename + + assay_df.to_csv(mod_filename, sep='\t', index=False) + print(f"Modified assay table saved as: {mod_filename}") + + except Exception as e: + report_failure_and_exit(str(e)) + + +if __name__ == "__main__": + main() \ No newline at end of file diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/update_curation_table.py b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/update_curation_table.py deleted file mode 100755 index a515def3d..000000000 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/bin/update_curation_table.py +++ /dev/null @@ -1,65 +0,0 @@ -#! /usr/bin/env python -""" Validation/Verification for raw reads in RNASeq Concensus Pipeline -""" -import argparse -from pathlib import Path - -import pandas as pd - -from dp_tools.core.post_processing import update_curation_tables -from dp_tools.core.loaders import load_data -from dp_tools.plugin_api import load_plugin - - -############################################################## -# Utility Functions To Handle Logging, Config and CLI Arguments -############################################################## -def _parse_args(): - """Parse command line args.""" - parser = argparse.ArgumentParser() - - parser.add_argument("--root-path", required=True, help="Root data path") - - parser.add_argument("--runsheet-path", required=True, help="Runsheet path") - - parser.add_argument("--plug-in-dir", required=True, help="Plugin path") - - parser.add_argument("--isa-path", required=True, help="ISA Archive Path") - - args = parser.parse_args() - return args - - -def main(root_dir: Path, runsheet_path: Path, plug_in_dir: Path, isa_path: Path): - plugin = load_plugin(Path(plug_in_dir)) - - ds = load_data( - config=plugin.config, - root_path=(root_dir), - runsheet_path=runsheet_path, - ) - - # inject ISA data asset - class ISA_ASSET: - path = isa_path - config = { - "resource categories": {"publish to repo": False} - } - - ds.dataset.data_assets["ISA Archive"] = ISA_ASSET() - isa_path = ds.dataset.data_assets["ISA Archive"].path - - update_curation_tables(ds.dataset, config=plugin.config) - - -if __name__ == "__main__": - import logging - - logging.basicConfig(level=logging.DEBUG) - log = logging.getLogger(__name__) - args = _parse_args() - main(Path(args.root_path), - runsheet_path=Path(args.runsheet_path), - plug_in_dir=args.plug_in_dir, - isa_path=args.isa_path - ) diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/config/default.config b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/config/default.config index b63c87727..10677552e 100644 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/config/default.config +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/config/default.config @@ -2,53 +2,32 @@ nextflow.enable.moduleBinaries = true params { - /* Here GLDS and OSD accession are defined. - Default behaviour is as follows: - - If accessions are not set, then either runsheet or an ISA Archive MUST be supplied - - If both accessions are set: - - If runsheet and ISA archive are left unset, then the ISA archive will be fetched from the GeneLab API and runsheet generated from the runsheet. - - If either runsheet or ISA archive are set, they will be used but the output directory and tags will reflect the appropriate accessions. This is useful when processing from the OSDR but OSDR metadata is not ready as is. - - If both runsheet and ISA archive are set, the workflow will halt. - - If only one accession is set, then the workflow will halt. - - */ - gldsAccession = "NOT_OSDR" // GeneLab Data Accession Number, e.g. GLDS-104 - osdAccession = "NOT_OSDR" // OSD Data Accession Number, e.g. OSD-367 - - // Catch case where only one is set - if (params.gldsAccession != "NOT_OSDR" && params.osdAccession == "NOT_OSDR") { - println "ERROR: GLDS accession set but OSD accession is not set. Please set both or neither." - System.exit(1) - } - if (params.gldsAccession == "NOT_OSDR" && params.osdAccession != "NOT_OSDR") { - println "ERROR: OSD accession set but GLDS accession is not set. Please set both or neither." - System.exit(1) - } - - resultsDir = (params.gldsAccession != "NOT_OSDR" && params.osdAccession != "NOT_OSDR") ? "./${params.gldsAccession}" : "." // the location for the output from the pipeline (also includes raw data and metadata) - /* Parameters that CAN be overwritten */ - runsheetPath = false - referenceStorePath = './References' // Path to custom references - biomart_attribute = false // Must be supplied if runsheet 'Array design REF' column doesn't indicate it - isaArchivePath = false // Alternative to fetching the ISA archive for an associated OSD/GLDS dataset - publish_dir_mode = "link" // method for creating publish directory. Default here for hardlink + accession = null + runsheetPath = "" // Path to a runsheet file, if not supplied, the workflow will attempt to generate one from the OSDR + isaArchivePath = "" // If a runsheet is not supplied, ISA.zip will be pulled from OSDR unless this is supplied + + api_url = "https://visualization.osdr.nasa.gov/biodata/api/v2/dataset/*/" // GeneLab API URL + dp_tools_plugin = null + + biomart_attribute = "" // Must be supplied if runsheet 'Array design REF' column doesn't indicate it + referenceStorePath = './References' // Path to custom annotation references + array_annot_path = "${projectDir}/../examples/annotations/design_info.csv" // // Probe annotation for non-Biomart supported microarrays + outdir = "${launchDir}" + publish_dir_mode = "link" // method for creating publish directory. Default here for hard link help = false // display help menu and exit workflow program /* Parameters that SHOULD NOT be overwritten */ - // For now, this particular is good for all organisms listed on the file. - annotation_file_path = "https://raw.githubusercontent.com/nasa/GeneLab_Data_Processing/GL_RefAnnotTable_1.0.0/GeneLab_Reference_Annotations/Pipeline_GL-DPPD-7110_Versions/GL-DPPD-7110/GL-DPPD-7110_annotations.csv" + annotation_file_path = "https://raw.githubusercontent.com/nasa/GeneLab_Data_Processing/GL_RefAnnotTable-A_1.1.0/GeneLab_Reference_Annotations/Pipeline_GL-DPPD-7110_Versions/GL-DPPD-7110-A/GL-DPPD-7110-A_annotations.csv" /* DEBUG related parameters, not likely useful in production */ skipVV = false // if true, VV will not be performed skipDE = false // if true, DE will not be performed - limit_biomart_query = false // if set to a value, that value is the maximum number of biomart probe IDs to query max_flag_code = 80 // Maximum flag value allowed, exceeding this value during V&V will cause the workflow to halt - } \ No newline at end of file diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/config/software/by_docker_image.config b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/config/software/by_docker_image.config index 6744407fd..b6012ed47 100644 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/config/software/by_docker_image.config +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/config/software/by_docker_image.config @@ -1,8 +1,10 @@ process { withName: 'PROCESS_AFFYMETRIX' { - container = "quay.io/j_81/gl_images:NF_AffyMP-A_1.0.0-RC7" + conda = "${projectDir}/envs/R_microarray_affymetrix.yaml" + container = "quay.io/nasa_genelab/gl-microarray:1.1.0" } - withName: 'RUNSHEET_FROM_GLDS|RUNSHEET_FROM_ISA|VV_AFFYMETRIX|GENERATE_MD5SUMS|UPDATE_ISA_TABLES|GENERATE_SOFTWARE_TABLE' { - container = "quay.io/j_81/dp_tools:1.3.4" + withName: 'GET_ACCESSIONS|FETCH_ISA|ISA_TO_RUNSHEET|VV_AFFYMETRIX|GENERATE_MD5SUMS|UPDATE_ASSAY_TABLE|GENERATE_SOFTWARE_TABLE|GENERATE_PROTOCOL' { + conda = "${projectDir}/envs/dp_tools.yaml" + container = "quay.io/nasa_genelab/dp_tools:1.3.8" } } diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/envs/R_microarray_affymetrix.yaml b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/envs/R_microarray_affymetrix.yaml new file mode 100644 index 000000000..f5874a607 --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/envs/R_microarray_affymetrix.yaml @@ -0,0 +1,19 @@ +name: R_microarray_affymetrix +channels: + - conda-forge + - bioconda + - nodefaults +dependencies: + - R-base=4.5.3 + - r-dplyr=1.2.0 + - r-ggplot2=4.0.2 + - r-glue=1.8.0 + - r-stringr=1.6.0 + - bioconductor-biocversion=3.22.0 + - bioconductor-limma=3.66.0 + - r-matrixstats=1.5.0 + - r-statmod=1.5.1 + - quarto=1.9.37 + - r-dt=0.34.0 + - r-downlit=0.4.5 + - r-xml2=1.6.0 diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/envs/dp_tools.yaml b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/envs/dp_tools.yaml new file mode 100644 index 000000000..8eb8ab2df --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/envs/dp_tools.yaml @@ -0,0 +1,11 @@ +name: dp_tools +channels: + - conda-forge + - bioconda + - nodefaults +dependencies: + - python=3.10 + - pip + - setuptools=81.0.0 + - pip: + - git+https://github.com/torres-alexis/dp_tools.git@v1.3.8 diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/main.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/main.nf index 3e7182721..6e8efcec9 100644 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/main.nf +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/main.nf @@ -1,132 +1,174 @@ nextflow.enable.dsl=2 -// color defs -c_back_bright_red = "\u001b[41;1m"; -c_bright_green = "\u001b[32;1m"; -c_blue = "\033[0;34m"; -c_reset = "\033[0m"; - -include { PARSE_ANNOTATION_TABLE } from './modules/PARSE_ANNOTATION_TABLE.nf' -include { VV_AFFYMETRIX } from './modules/VV_AFFYMETRIX.nf' -include { PROCESS_AFFYMETRIX } from './modules/PROCESS_AFFYMETRIX.nf' -include { RUNSHEET_FROM_GLDS } from './modules/RUNSHEET_FROM_GLDS.nf' -include { RUNSHEET_FROM_ISA } from './modules/RUNSHEET_FROM_ISA.nf' -include { GENERATE_SOFTWARE_TABLE } from './modules/GENERATE_SOFTWARE_TABLE' -include { DUMP_META } from './modules/DUMP_META' - -/************************************************** -* HELP MENU ************************************** -**************************************************/ -if (params.help) { - println("┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅") - println("┇ Affymetrix Microarray Pipeline: $workflow.manifest.version ┇") - println("┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅┅") - println("Usage example 1: Processing GLDS datasets using genome fasta and gtf from Ensembl") - println(" > nextflow run ./main.nf --osdAccession OSD-266 --gldsAccession GLDS-266") - println() - println("Usage example 2: Processing Other datasets") - println(" Note: This requires a user-created runsheet.") - println(" > nextflow run ./main.nf --runsheetPath ") - println() - println("arguments:") - println(" --help show this help message and exit") - println(" --osdAccession OSD-000") - println(" the OSD accession id to process through the Affymetrix Microarray Pipeline.") - println(" --gldsAccession GLDS-000") - println(" the GLDS accession id to process through the Affymetrix Microarray Pipeline.") - println(" --runsheetPath Use a local runsheet instead one automatically generated from a GLDS ISA archive.") - println(" --skipVV Skip automated V&V. Default: false") - println(" --skipDE Skip DE. Default: false") - println(" --resultsDir Directory to save staged raw files and processed files. Default: ") - exit 0 - } - -println "PARAMS: $params" -println "\n" - -/************************************************** -* CHECK REQUIRED PARAMS AND LOAD ***************** -**************************************************/ -println("Resolved output directory: ${ params.resultsDir }") - -/************************************************** -* WORKFLOW SPECIFIC PRINTOUTS ******************** -**************************************************/ + +include { paramsHelp } from 'plugin/nf-schema' +include { validateParameters } from 'plugin/nf-schema' +include { paramsSummaryLog } from 'plugin/nf-schema' + +include { STAGE_ANALYSIS } from './subworkflows/stage_analysis.nf' +include { PARSE_ANNOTATION_TABLE } from './modules/parse_annotation_table.nf' +include { VV_AFFYMETRIX } from './modules/vv_affymetrix.nf' +include { PROCESS_AFFYMETRIX } from './modules/process_affymetrix.nf' +include { GENERATE_SOFTWARE_TABLE } from './modules/generate_software_table.nf' +include { DUMP_META } from './modules/dump_meta.nf' +include { GENERATE_PROTOCOL } from './modules/generate_protocol.nf' + +ch_dp_tools_plugin = params.dp_tools_plugin ? + channel.value(file(params.dp_tools_plugin)) + : params.skipDE ? channel.value(file("$projectDir/bin/dp_tools__affymetrix_skipDE")) : channel.value(file("$projectDir/bin/dp_tools__affymetrix")) +ch_isa_archive_path = params.isaArchivePath ? file(params.isaArchivePath) : null +ch_runsheet = params.runsheetPath ? channel.fromPath(params.runsheetPath) : null + +ch_outdir = params.outdir ? channel.fromPath(params.outdir, checkIfExists: true) : null workflow { main: - if ( !params.runsheetPath && !params.isaArchivePath) { - RUNSHEET_FROM_GLDS( - params.osdAccession, - params.gldsAccession, - "${ projectDir }/bin/dp_tools__affymetrix" // dp_tools plugin - ) - RUNSHEET_FROM_GLDS.out.runsheet | set{ ch_runsheet } - } else if ( !params.runsheetPath && params.isaArchivePath ) { - RUNSHEET_FROM_ISA( - params.osdAccession, - params.gldsAccession, - params.isaArchivePath, - "${ projectDir }/bin/dp_tools__affymetrix" // dp_tools plugin - ) - RUNSHEET_FROM_ISA.out.runsheet | set{ ch_runsheet } - } else if ( params.runsheetPath && !params.isaArchivePath ) { - ch_runsheet = channel.fromPath( params.runsheetPath ) - } else if ( params.runsheetPath && params.isaArchivePath ) { - System.err.println("Error: User supplied both runsheetPath and isaArchivePath. Only one or neither is allowed to be supplied!") // Print error message to System.err - System.exit(1) // Exit with error code 1 - } + // color defs + c_back_bright_red = "\u001b[41;1m"; + c_bright_green = "\u001b[32;1m"; + c_blue = "\033[0;34m"; + c_reset = "\033[0m"; - ch_runsheet | splitCsv(header: true) | first | view | set{ ch_meta } + /************************************************** + * HELP MENU ************************************** + **************************************************/ + if (params.help) { + before_text = """ +************************************************ +* Microarray Affymetrix Pipeline: ${workflow.manifest.version} * +************************************************ - PARSE_ANNOTATION_TABLE( - params.annotation_file_path, - ch_meta | map { it.organism } - ) +Usage example 1: Processing OSDR datasets + > nextflow run ./main.nf --osdAccession OSD-266 --gldsAccession GLDS-266 + +Usage example 2: Processing Other datasets (requires a user-created runsheet) + > nextflow run ./main.nf --runsheetPath + + +""" + log.info paramsHelp( + beforeText: before_text, + afterText: "For more information, please see the README.md file in the workflow code directory.", + fullHelp: true) + exit(0) + } + + // validate parameters and print parameter summary log (includes only parameters that are not set to default values) + validateParameters(cast_cli_params: true) + log.info paramsSummaryLog(workflow) + + // --------------------- Sanity Checks ------------------------------------- // + // Test input requirement + if (!params.accession && !params.runsheetPath){ + error("""${c_back_bright_red}INPUT ERROR! + Please supply either an accession (OSD or Genelab number) or an input CSV file + by passing either to the --accession or --runsheetPath parameter, respectively. + ${c_reset}""") + } + + // Test ISA archive and accession + if (params.isaArchivePath && !params.accession) { + error """${c_back_bright_red}INPUT ERROR! + --isaArchivePath requires --accession to resolve OSD/GLDS accessions + for the ISA-to-runsheet conversion.${c_reset}""" + } + + // Stage analysis setup (directory structure, inputs, and raw reads) + STAGE_ANALYSIS( + ch_outdir, + ch_dp_tools_plugin, + params.accession, + ch_isa_archive_path, + ch_runsheet, + params.api_url + ) + ch_outdir = STAGE_ANALYSIS.out.ch_outdir + samples = STAGE_ANALYSIS.out.samples + array_data_files = STAGE_ANALYSIS.out.array_data_files | map { meta, f -> f } | collect + runsheet_path = STAGE_ANALYSIS.out.runsheet_path + isa_archive = STAGE_ANALYSIS.out.isa_archive + osd_accession = STAGE_ANALYSIS.out.osd_accession + glds_accession = STAGE_ANALYSIS.out.glds_accession + dp_tools_version = STAGE_ANALYSIS.out.dp_tools_version + + // Get dataset-wide metadata + samples | first + | map { meta, reads -> meta } + | set { ch_meta } + + ch_meta | map { meta -> meta.organism_sci } + | set { organism_sci } + + + PARSE_ANNOTATION_TABLE(params.annotation_file_path, organism_sci) PROCESS_AFFYMETRIX( + ch_outdir, channel.fromPath( "${ projectDir }/bin/Affymetrix.qmd" ), - ch_runsheet, + runsheet_path, + array_data_files, PARSE_ANNOTATION_TABLE.out.annotations_db_url, PARSE_ANNOTATION_TABLE.out.reference_version_and_source, - params.limit_biomart_query, + channel.fromPath( params.referenceStorePath ), + channel.fromPath( params.array_annot_path ), params.skipDE ) VV_AFFYMETRIX( - ch_runsheet, + ch_outdir, + runsheet_path, PROCESS_AFFYMETRIX.out.de, params.skipVV, - "${ projectDir }/bin/${ params.skipDE ? 'dp_tools__affymetrix_skipDE' : 'dp_tools__affymetrix' }" // dp_tools plugin + ch_dp_tools_plugin ) // Software Version Capturing - nf_version = "- name: nextflow\n ".concat( -""" - version: ${nextflow.version} - homepage: https://www.nextflow.io - workflow task: N/A -""") - ch_software_versions = Channel.value(nf_version) - PROCESS_AFFYMETRIX.out.versions | map{ it -> it.text } | mix(ch_software_versions) | set{ch_software_versions} - VV_AFFYMETRIX.out.versions | map{ it -> it.text } | mix(ch_software_versions) | set{ch_software_versions} + ch_software_versions = channel.empty() + nf_version = """\ + - name: nextflow + version: ${nextflow.version} + homepage: https://www.nextflow.io + workflow task: N/A + """.stripIndent() + ch_nextflow_version = channel.value(nf_version) + ch_process_AFFYMETRIX = PROCESS_AFFYMETRIX.out.versions | map{ it -> it.text } + ch_dp_tools_version = params.skipVV + ? ( dp_tools_version ? dp_tools_version.map { it -> it.text} : channel.empty() ) + : ( VV_AFFYMETRIX.out.versions | map{ it -> it.text } ) + + ch_software_versions = ch_software_versions + | mix(ch_process_AFFYMETRIX) + | mix(ch_dp_tools_version) + | mix(ch_nextflow_version) GENERATE_SOFTWARE_TABLE( + ch_outdir, ch_software_versions | unique | collectFile(newLine: true, sort: true, cache: false), - ch_runsheet | splitCsv(header: true, quote: '"') | first | map{ row -> row['Array Data File Name'] }, + runsheet_path | splitCsv(header: true, quote: '"') | first | map{ row -> row['Array Data File Name'] }, params.skipDE ) // export meta for post processing usage - ch_meta | DUMP_META -} + DUMP_META(ch_outdir, ch_meta) + + GENERATE_PROTOCOL( + ch_outdir, + ch_meta, + ch_software_versions | unique | collectFile(newLine: true, sort: true, cache: false), + PARSE_ANNOTATION_TABLE.out.reference_version_and_source, + PARSE_ANNOTATION_TABLE.out.bioconductor_annotations, + PARSE_ANNOTATION_TABLE.out.annotations_db_info_url, + params.skipDE + ) -workflow.onComplete { - println "${c_bright_green}Pipeline completed at: $workflow.complete" - println "Execution status: ${ workflow.success ? 'OK' : 'failed' }" - if ( workflow.success ) { - println "Raw and Processed data location: ${ params.resultsDir }" - println "V&V logs location: ${ params.resultsDir }/VV_Logs" - println "Pipeline tracing/visualization files location: ${ params.resultsDir }/Resource_Usage${c_reset}" + ch_outdir.subscribe { outdir_value = it } + workflow.onComplete = { + println "${c_bright_green}Pipeline completed at: $workflow.complete" + println "Execution status: ${ workflow.success ? 'OK' : 'failed' }" + if ( workflow.success ) { + println "Raw and Processed data location: ${ outdir_value }" + println "V&V logs location: ${ outdir_value }/VV_Logs" + println "Pipeline tracing/visualization files location: ${ outdir_value }/Resource_Usage${c_reset}" + } } } diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/GENERATE_MD5SUMS.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/GENERATE_MD5SUMS.nf deleted file mode 100644 index deae04413..000000000 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/GENERATE_MD5SUMS.nf +++ /dev/null @@ -1,22 +0,0 @@ -process GENERATE_MD5SUMS { - // Generates tabular data indicating genelab standard publishing files, md5sum generation, and tool version table formatting - tag "${ params.gldsAccession }" - publishDir "${ params.resultsDir }/GeneLab", - mode: params.publish_dir_mode - - input: - path(data_dir) - path(runsheet) - path(dp_tools__affymetrix) - - output: - path("*md5sum*") - path("Missing_md5sum_files_GLmicroarray.txt"), optional: true - - script: - """ - generate_md5sum_files.py --root-path ${ data_dir } \\ - --runsheet-path ${ runsheet } \\ - --plug-in-dir ${ dp_tools__affymetrix } - """ -} \ No newline at end of file diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/PARSE_ANNOTATION_TABLE.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/PARSE_ANNOTATION_TABLE.nf deleted file mode 100644 index a6a392810..000000000 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/PARSE_ANNOTATION_TABLE.nf +++ /dev/null @@ -1,38 +0,0 @@ -process PARSE_ANNOTATION_TABLE { - // Extracts data from this kind of table: - // https://github.com/nasa/GeneLab_Data_Processing/blob/master/GeneLab_Reference_Annotations/Pipeline_GL-DPPD-7110_Versions/GL-DPPD-7110/GL-DPPD-7110_annotations.csv - - input: - val(annotations_csv_url_string) - val(organism_sci) - - output: - val(annotations_db_url), emit: annotations_db_url - tuple val(ensemblVersion), val(ensemblSource), emit: reference_version_and_source - - exec: - def organisms = [:] - println "Fetching table from ${annotations_csv_url_string}" - - // download data to memory - annotations_csv_url_string.toURL().splitEachLine(",") {fields -> - organisms[fields[1]] = fields - } - // extract required fields - organism_key = organism_sci.capitalize().replace("_"," ") - // fasta_url = organisms[organism_key][5] - // gtf_url = organisms[organism_key][6] - annotations_db_url = organisms[organism_key][9] - ensemblVersion = organisms[organism_key][3] - ensemblSource = organisms[organism_key][4] - - println "PARSE_ANNOTATION_TABLE:" - println "Values parsed for '${organism_key}' using process:" - println "--------------------------------------------------" - // println "- fasta_url: ${fasta_url}" - // println "- gtf_url: ${gtf_url}" - println "- annotations_db_url: ${annotations_db_url}" - println "- ensemblVersion: ${ensemblVersion}" - println "- ensemblSource: ${ensemblSource}" - println "--------------------------------------------------" -} diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/POST_PROCESSING/GENERATE_PROTOCOL/main.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/POST_PROCESSING/GENERATE_PROTOCOL/main.nf deleted file mode 100644 index 32a0544f3..000000000 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/POST_PROCESSING/GENERATE_PROTOCOL/main.nf +++ /dev/null @@ -1,18 +0,0 @@ -process GENERATE_PROTOCOL { - tag "${ params.gldsAccession }" - publishDir "${ params.resultsDir }/GeneLab", - mode: params.publish_dir_mode - - input: - path("software_versions_GLmicroarray.md") - path("meta.sh") - val(skipDE) - - output: - path("PROTOCOL_GLmicroarray.txt") - - script: - """ - generate_protocol.sh $workflow.manifest.version $skipDE - """ -} \ No newline at end of file diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/POST_PROCESSING/GENERATE_PROTOCOL/resources/usr/bin/generate_protocol.sh b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/POST_PROCESSING/GENERATE_PROTOCOL/resources/usr/bin/generate_protocol.sh deleted file mode 100755 index 5b6f2357b..000000000 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/POST_PROCESSING/GENERATE_PROTOCOL/resources/usr/bin/generate_protocol.sh +++ /dev/null @@ -1,86 +0,0 @@ -#!/bin/bash -set -u - -software_versions_file="software_versions_GLmicroarray.md" - -# Read the markdown table -while read -r line; do - # Extract program, version, and link - program=$(echo "$line" | awk -F'|' '{gsub(/^[[:blank:]]+|[[:blank:]]+$/,"",$1); print $1}') - version=$(echo "$line" | awk -F'|' '{gsub(/^[[:blank:]]+|[[:blank:]]+$/,"",$2); print $2}') - - # Skip the header row and rows without version information - if [[ $program != "Program" && $version != "Version" && ! -z $version ]]; then - # Replace invalid characters in program name with underscores - sanitized_program=$(echo "$program" | tr -cd '[:alnum:]_') - - # Create environment variable name - env_var_name="${sanitized_program}_VERSION" - - # Set the environment variable - export "$env_var_name=$version" - fi -done < <(sed -n '/|/p' "$software_versions_file" | sed 's/^ *|//;s/|$//') - -# Print the extracted versions -env | grep "_VERSION" - -# Determine mapped sections -source meta.sh - -# List of organisms -organism_list=("Homo sapiens" "Mus musculus" "Rattus norvegicus" "Drosophila melanogaster" "Caenorhabditis elegans" "Danio rerio" "Saccharomyces cerevisiae") - -# Check the value of 'organism' variable and set 'GENE_MAPPING_STEP' accordingly -if [[ $organism == "Arabidopsis thaliana" ]]; then - GENE_MAPPING_STEP="Ensembl gene ID mappings were retrieved for each probeset using the Plants Ensembl database ftp server (plants.ensembl.org, release 54)." -elif [[ $organism == "Escherichia coli" ]]; then - GENE_MAPPING_STEP="Gene annotations were retrieved for each probeset from ThermoFisher (https://www.thermofisher.com/order/catalog/product/sec/assets?url=TFS-Assets/LSG/Support-Files/E_coli_2-na36-annot-csv.zip, created March 2016, accessed June 2024)." -elif [[ $organism == "Pseudomonas aeruginosa" ]]; then - GENE_MAPPING_STEP="Gene annotations were retrieved for each probeset from ThermoFisher (https://www.thermofisher.com/order/catalog/product/sec/assets?url=TFS-Assets/LSG/Support-Files/Pae_G1a-na36-annot-csv.zip, created March 2016, accessed June 2024)." -elif [[ " ${organism_list[*]} " == *"${organism//\"/}"* ]]; then - GENE_MAPPING_STEP="Ensembl gene ID mappings were retrieved for each probeset using biomaRt (version ${biomaRt_VERSION}), Ensembl database (ensembl.org, release 107)." -else - GENE_MAPPING_STEP="TBD" -fi - -# Check the value of 'organism' variable and set 'GENE_MAPPING_STEP' accordingly -if [[ $organism == "Arabidopsis thaliana" ]]; then - GENE_ANNOTATION_DB="org.At.tair.db" -elif [[ $organism == "Homo sapiens" ]]; then - GENE_ANNOTATION_DB="org.Hs.eg.db" -elif [[ $organism == "Mus musculus" ]]; then - GENE_ANNOTATION_DB="org.Mm.eg.db" -elif [[ $organism == "Rattus norvegicus" ]]; then - GENE_ANNOTATION_DB="org.Rn.eg.db" -elif [[ $organism == "Drosophila melanogaster" ]]; then - GENE_ANNOTATION_DB="org.Dm.eg.db" -elif [[ $organism == "Caenorhabditis elegans" ]]; then - GENE_ANNOTATION_DB="org.Ce.eg.db" -elif [[ $organism == "Danio rerio" ]]; then - GENE_ANNOTATION_DB="org.Dr.eg.db" -elif [[ $organism == "Saccharomyces cerevisiae" ]]; then - GENE_ANNOTATION_DB="org.Sc.sgd.db" -else - GENE_ANNOTATION_DB="TBD" -fi - -# Check if DGE was performed -if $2; then - DE_STEP="" -else - DE_STEP="Differential expression analysis was performed in R (version ${R_VERSION}) using limma (version ${limma_VERSION}); all groups were compared pairwise for each probeset to generate a moderated t-statistic and associated p- and adjusted p-value." -fi - -# Gene annotations -if [[ $organism == "Escherichia coli" || $organism == "Pseudomonas aeruginosa" ]]; then - ANNOT_STEP="" # Already covered in GENE_MAPPING_STEP -else - ANNOT_STEP="Gene annotations were assigned for every probeset that mapped to exactly one Ensembl gene ID using the custom annotation tables generated in-house as detailed in GL-DPPD-7110 (https://github.com/nasa/GeneLab_Data_Processing/blob/GL_RefAnnotTable_1.0.0/GeneLab_Reference_Annotations/Pipeline_GL-DPPD-7110_Versions/GL-DPPD-7110/GL-DPPD-7110.md), with STRINGdb (version 2.8.4), PANTHER.db (version 1.0.11), and ${GENE_ANNOTATION_DB} (version 3.15.0)." -fi - -# Read the template file -template="Data were processed as described in GL-DPPD-7114 (https://github.com/nasa/GeneLab_Data_Processing/blob/master/Microarray/Affymetrix/Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114.md) using NF_MAAffymetrix version $1 (https://github.com/nasa/GeneLab_Data_Processing/tree/NF_MAAffymetrix_$1/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix). In short, a RunSheet containing raw data file location and processing metadata from the study's *ISA.zip file was generated using dp_tools (version ${dp_tools_VERSION}). The raw array data files were loaded into R (version ${R_VERSION}) using oligo (version ${oligo_VERSION}). Raw data quality assurance density plot, pseudo images, MA plots, and boxplots were generated using oligo (version ${oligo_VERSION}). The raw probe level intensity data was background corrected and normalized across arrays via the oligo (version ${oligo_VERSION}) quantile method. Normalized probe level data quality assurance density plot, pseudo images, MA plots, and boxplots were generated using oligo (version ${oligo_VERSION}). Normalized probe level data was summarized to the probeset level using the oligo (version ${oligo_VERSION}) RMA method. ${GENE_MAPPING_STEP} ${DE_STEP} ${ANNOT_STEP}" - -# Output the filled template -echo "$template" > PROTOCOL_GLmicroarray.txt \ No newline at end of file diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/RUNSHEET_FROM_GLDS.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/RUNSHEET_FROM_GLDS.nf deleted file mode 100644 index 10777f10f..000000000 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/RUNSHEET_FROM_GLDS.nf +++ /dev/null @@ -1,28 +0,0 @@ -process RUNSHEET_FROM_GLDS { - // Downloads isa Archive and creates runsheet using GeneLab API - tag "${ gldsAccession }" - publishDir "${ params.resultsDir }/Metadata", - pattern: "*.zip", - mode: params.publish_dir_mode - - input: - val(osdAccession) - val(gldsAccession) - path("dp_tools__affymetrix") - - output: - path("${ osdAccession }_microarray_v?_runsheet.csv"), emit: runsheet - path("*.zip"), emit: isaArchive - - script: - def injects = params.biomart_attribute ? "--inject biomart_attribute='${ params.biomart_attribute }'" : '' - """ - - dpt-get-isa-archive --accession ${ osdAccession } - ls dp_tools__affymetrix - - dpt-isa-to-runsheet --accession ${ osdAccession } \ - --plugin-dir dp_tools__affymetrix \ - --isa-archive *.zip ${ injects } - """ -} \ No newline at end of file diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/RUNSHEET_FROM_ISA.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/RUNSHEET_FROM_ISA.nf deleted file mode 100644 index c00a097c8..000000000 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/RUNSHEET_FROM_ISA.nf +++ /dev/null @@ -1,23 +0,0 @@ -process RUNSHEET_FROM_ISA { - // Generates Runsheet using a path to an ISA archive - tag "${ gldsAccession }" - publishDir "${ params.resultsDir }/Metadata", - pattern: "*.zip", - mode: params.publish_dir_mode - - input: - val(osdAccession) - val(gldsAccession) - path(isaArchive) - path("dp_tools__agilent_1_channel") - - output: - path("${ osdAccession }_microarray_v?_runsheet.csv"), emit: runsheet - - script: - """ - dpt-isa-to-runsheet --accession ${ osdAccession } \ - --plugin-dir dp_tools__agilent_1_channel \ - --isa-archive ${ isaArchive } - """ -} \ No newline at end of file diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/UPDATE_ISA_TABLES.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/UPDATE_ISA_TABLES.nf deleted file mode 100644 index 441cb9f6b..000000000 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/UPDATE_ISA_TABLES.nf +++ /dev/null @@ -1,25 +0,0 @@ -process UPDATE_ISA_TABLES { - // Generates tabular data indicating genelab standard publishing files, md5sum generation, and tool version table formatting - tag "${ params.gldsAccession }" - publishDir "${ params.resultsDir }/GeneLab", - mode: params.publish_dir_mode - - input: - path(data_dir) - path(runsheet) - path(dp_tools__affymetrix) - - output: - path("updated_curation_tables") // directory containing extended ISA tables - - script: - """ - update_curation_table.py --root-path ${ data_dir } \\ - --runsheet-path ${ runsheet } \\ - --plug-in-dir ${ dp_tools__affymetrix } \\ - --isa-path ${ data_dir }/Metadata/*ISA*.zip - - # Update assay table with gldsAccession - sed -i 's/${ params.osdAccession }/${ params.gldsAccession }/g' updated_curation_tables/a*.txt - """ -} \ No newline at end of file diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/copy_array_data_files.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/copy_array_data_files.nf new file mode 100755 index 000000000..817222e72 --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/copy_array_data_files.nf @@ -0,0 +1,23 @@ +// Nextflow's path staging fetches http(s)/ftp URLs (or copies local paths) automatically; +// this process just normalizes the result to the manufacturer's expected filename, +// decompressing if the source was gzipped. +process COPY_ARRAY_DATA_FILES { + tag "${ meta.id }" + + input: + tuple val(meta), path(array_data_file, stageAs: "staged_input") + + output: + tuple val(meta), path("${meta.file_name}") + + script: + if (meta.is_gz) { + """ + gunzip -c staged_input > "${meta.file_name}" + """ + } else { + """ + cp -P staged_input "${meta.file_name}" + """ + } +} diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/DUMP_META/main.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/dump_meta.nf similarity index 77% rename from Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/DUMP_META/main.nf rename to Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/dump_meta.nf index 2e0381a83..ad1c7346a 100644 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/DUMP_META/main.nf +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/dump_meta.nf @@ -1,8 +1,9 @@ process DUMP_META { - publishDir "${ params.resultsDir }/GeneLab", + publishDir "${ publishdir }/GeneLab", mode: params.publish_dir_mode input: + val(publishdir) val(meta) output: diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/fetch_isa.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/fetch_isa.nf new file mode 100755 index 000000000..73c82d3e0 --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/fetch_isa.nf @@ -0,0 +1,19 @@ +process FETCH_ISA { + tag "${osd_accession}_${glds_accession}" + + publishDir "${publishdir}/Metadata", + mode: params.publish_dir_mode + + input: + val(publishdir) + val(osd_accession) + val(glds_accession) + + output: + path "*.zip", emit: isa_archive + + script: + """ + fetch_isa.py --osd ${osd_accession} --outdir . + """ +} diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/generate_md5sums.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/generate_md5sums.nf new file mode 100644 index 000000000..21558197a --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/generate_md5sums.nf @@ -0,0 +1,17 @@ +process GENERATE_MD5SUMS { + // Generates tabular data indicating genelab standard publishing files, md5sum generation, and tool version table formatting + publishDir "${ data_dir }/GeneLab", + mode: params.publish_dir_mode + + input: + path(data_dir) + val(done_token) // ensures process runs after purging processing_info.txt + + output: + path("processed_md5sum_GLmicroarray.tsv"), emit: processed_md5sum + + script: + """ + generate_md5sum_files.py --outdir ${data_dir} --workflow_version ${workflow.manifest.version} + """ +} \ No newline at end of file diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/generate_protocol.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/generate_protocol.nf new file mode 100644 index 000000000..f5a4e6f18 --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/generate_protocol.nf @@ -0,0 +1,39 @@ +process GENERATE_PROTOCOL { + publishDir "${ publishdir }/GeneLab", + mode: params.publish_dir_mode, + pattern: "*.txt" + + input: + val(publishdir) + val(ch_meta) + path(software_versions_yaml) + tuple val(ensemblVersion), val(ensemblSource) + val(bioconductor_annotations) + path(annotations_db_info) + val(skipDE) + + output: + path("protocol_GLmicroarray.txt") + + script: + def skipDGE = skipDE ? "--skip-DGE" : '' + def custom_annot_file = params.array_annot_path ? "--custom_annot_design ${params.array_annot_path}" : '' + def bioconductor_annot_file = bioconductor_annotations ? "--bioconductor_annotations ${bioconductor_annotations}" : '' + def annot_db_info_file = annotations_db_info ? "--annotations_db_info ${annotations_db_info}" : '' + + """ + generate_protocol.py \ + --outdir . \ + --software_table ${software_versions_yaml} \ + --assay_suffix "_GLmicroarray" \ + --workflow_version ${workflow.manifest.version} \ + --organism "${ch_meta.organism_sci}" \ + --reference_source ${ensemblSource} \ + --reference_version ${ensemblVersion} \ + --biomart_attribute "${ch_meta.biomart_id}" \ + $bioconductor_annot_file \ + $annot_db_info_file \ + $custom_annot_file \ + $skipDGE + """ +} \ No newline at end of file diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/GENERATE_SOFTWARE_TABLE/main.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/generate_software_table.nf similarity index 85% rename from Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/GENERATE_SOFTWARE_TABLE/main.nf rename to Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/generate_software_table.nf index 6d260d303..03b4fcd8f 100644 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/GENERATE_SOFTWARE_TABLE/main.nf +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/generate_software_table.nf @@ -1,9 +1,10 @@ process GENERATE_SOFTWARE_TABLE { - publishDir "${ params.resultsDir }/GeneLab", + publishDir "${ publishdir }/GeneLab", pattern: "software_versions_GLmicroarray.md", mode: params.publish_dir_mode input: + val(publishdir) path("software_versions.yaml") val(filename) val(skipDE) diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/get_accessions.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/get_accessions.nf new file mode 100755 index 000000000..78c417edc --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/get_accessions.nf @@ -0,0 +1,20 @@ +// Take the accession in the format 'OSD-###' or 'GLDS-###' +// Query the API to Find all GLDS accessions associated with the OSD accessions +// Create a dictionary with keys 'osd_accession' values 'glds_accession'. +// Then return the osd_accession and glds_accession. + +process GET_ACCESSIONS { + tag "${accession}" + + input: + val accession + val api_url + + output: + path "accessions.txt", emit: accessions_txt + + script: + """ + get_accessions.py --accession "${accession}" --api_url "${api_url}" > accessions.txt + """ +} \ No newline at end of file diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/isa_to_runsheet.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/isa_to_runsheet.nf new file mode 100755 index 000000000..9bf33a781 --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/isa_to_runsheet.nf @@ -0,0 +1,44 @@ +process ISA_TO_RUNSHEET { + tag "${osd_accession}_${glds_accession}" + + publishDir "${publishdir}/Metadata", + mode: params.publish_dir_mode, + pattern: "*.csv" + + publishDir "${publishdir}/Metadata", + mode: params.publish_dir_mode, + pattern: "isa_archive/*", + saveAs: { filename -> + if (filename.startsWith("isa_archive/")) return filename.replace("isa_archive/", "") + else return filename + } + + input: + val(publishdir) + val(osd_accession) + val(glds_accession) + path(isa_archive) + path(dp_tools_plugin) + + output: + path("*.csv"), emit: runsheet + path("isa_archive/${isa_archive}") + path("versions.yml"), emit: version + + script: + """ + dpt-isa-to-runsheet --accession ${osd_accession} --isa-archive ${isa_archive} --plugin-dir ${dp_tools_plugin} + + # Copy the ISA archive to the output directory + mkdir -p isa_archive + cp ${isa_archive} isa_archive/ + + # Export dp_tools version + cat >> versions.yml < + def values = line.strip().split(",") + [headers, values].transpose().collectEntries() + } + + def organism_key = organism_sci.capitalize().replace("_"," ") + + def organism_record = records.find { rec -> rec['species'] == organism_key } + if (organism_record == null) { + throw new Exception("Organism '${organism_key}' not found in annotation table at ${annotations_csv_url_string}") + } else { + annotations_db_url = organism_record['genelab_annots_link'] + ensemblVersion = organism_record['ensemblVersion'] + ensemblSource = organism_record['ref_source'] + bioconductor_annotations = organism_record['bioconductor_annotations'] + annotations_db_info_url = organism_record['genelab_annots_info_link'] + + // Convert figshare ndownloader URL to API endpoint + if (annotations_db_url != null && annotations_db_url.contains('figshare.com/ndownloader/files/')) { + file_id = (annotations_db_url =~ /.*\/files\/([a-zA-Z0-9]+).*/)[0][1] + annotations_db_url = "https://api.figshare.com/v2/file/download/${file_id}" + } + + // Convert figshare ndownloader URL to API endpoint + if (annotations_db_info_url != null && annotations_db_info_url.contains('figshare.com/ndownloader/files/')) { + file_id = (annotations_db_info_url =~ /.*\/files\/([a-zA-Z0-9]+).*/)[0][1] + annotations_db_info_url = "https://api.figshare.com/v2/file/download/${file_id}" + } + + println "PARSE_ANNOTATION_TABLE:" + println "Values parsed for '${organism_key}' using process:" + println "--------------------------------------------------" + println "- annotations_db_url: ${annotations_db_url}" + println "- annotations_db_info_url: ${annotations_db_info_url}" + println "- ensemblVersion: ${ensemblVersion}" + println "- ensemblSource: ${ensemblSource}" + println "--------------------------------------------------" + } +} diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/PROCESS_AFFYMETRIX.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/process_affymetrix.nf similarity index 73% rename from Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/PROCESS_AFFYMETRIX.nf rename to Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/process_affymetrix.nf index 6f6e33013..5059aa465 100644 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/PROCESS_AFFYMETRIX.nf +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/process_affymetrix.nf @@ -1,15 +1,18 @@ process PROCESS_AFFYMETRIX { - publishDir "${ params.resultsDir }/GeneLab", + publishDir "${ publishdir }/GeneLab", pattern: "NF_MAAffymetrix_v${workflow.manifest.version}_GLmicroarray.html", mode: params.publish_dir_mode stageInMode 'copy' input: + val(publishdir) path(qmd) // quarto qmd file to render path(runsheet_csv) // runsheet to supply as parameter + path(array_data_files) // staged, locally-named, decompressed raw array data files path(annotation_file_path) tuple val(ensemblVersion), val(ensemblSource) - val(limit_biomart_query) // DEBUG option, limits biomart queries to the number specified if not set to false + path(referenceStorePath) // path to custom annotation references + path(array_annot_path) // path to custom array design info file val(skipDE) // whether to skip DE output: @@ -22,17 +25,17 @@ process PROCESS_AFFYMETRIX { path("versions.yml"), emit: versions script: - def limit_biomart_query_parameter = limit_biomart_query ? "-P DEBUG_limit_biomart_query:${limit_biomart_query}" : '' def run_DE = skipDE ? "-P run_DE:'false'" : '' """ export HOME=\$PWD; quarto render \$PWD/${qmd} \ + -P 'workflow_version:${workflow.manifest.version}' \ -P 'runsheet:${runsheet_csv}' \ -P 'annotation_file_path:${annotation_file_path}' \ -P 'ensembl_version:${ensemblVersion}' \ - -P 'local_annotation_dir:${params.referenceStorePath}' \ - ${limit_biomart_query_parameter} \ + -P 'local_annotation_dir:${referenceStorePath}' \ + -P 'array_annot_path:${array_annot_path}' \ ${run_DE} # Rename report diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/purge_processing_info.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/purge_processing_info.nf new file mode 100755 index 000000000..586d816d5 --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/purge_processing_info.nf @@ -0,0 +1,18 @@ +process PURGE_PROCESSING_INFO { + + publishDir "${ data_dir }/GeneLab", + mode: params.publish_dir_mode + + input: + path(data_dir) + path(processing_info, stageAs: "unpurged_processing_info_GLmicroarray.txt") + + output: + path("nextflow_processing_info_GLmicroarray.txt"), emit: purged_processing_info + + script: + """ + array_annot_basename=\$(basename '${params.array_annot_path}') + sed "s|[^ ]*/\${array_annot_basename}|\${array_annot_basename}|g" ${processing_info} > nextflow_processing_info_GLmicroarray.txt + """ +} \ No newline at end of file diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/update_assay_table.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/update_assay_table.nf new file mode 100644 index 000000000..beed1ccea --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/update_assay_table.nf @@ -0,0 +1,21 @@ +process UPDATE_ASSAY_TABLE { + // Generates tabular data indicating genelab standard publishing files, md5sum generation, and tool version table formatting + publishDir "${ data_dir }/GeneLab/updated_curation_tables", + mode: params.publish_dir_mode + + input: + path(data_dir) + path(runsheet) + path(isa_archive) + val(glds_accession) + + output: + path("a_*.txt"), emit: updated_assay_table + + script: + """ + update_assay_table.py --runsheet ${ runsheet } \ + --glds_accession ${ glds_accession } \ + --isa_zip ${ isa_archive } + """ +} \ No newline at end of file diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/VV_AFFYMETRIX.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/vv_affymetrix.nf similarity index 85% rename from Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/VV_AFFYMETRIX.nf rename to Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/vv_affymetrix.nf index fb56f022d..a551a3711 100644 --- a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/VV_AFFYMETRIX.nf +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/modules/vv_affymetrix.nf @@ -1,26 +1,27 @@ process VV_AFFYMETRIX { // Log publishing - publishDir "${ params.resultsDir }", + publishDir "${ publishdir }", pattern: "VV_report_GLmicroarray.tsv.MANUAL_CHECKS_PENDING" , mode: params.publish_dir_mode, saveAs: { "VV_Logs/VV_log_${ task.process.replace(":","-") }_GLmicroarray.tsv.MANUAL_CHECKS_PENDING" } // V&V'ed data publishing - publishDir "${ params.resultsDir }", + publishDir "${ publishdir }", pattern: '00-RawData/**', mode: params.publish_dir_mode - publishDir "${ params.resultsDir }", + publishDir "${ publishdir }", pattern: '01-oligo_NormExp/**', mode: params.publish_dir_mode - publishDir "${ params.resultsDir }", + publishDir "${ publishdir }", pattern: '02-limma_DGE/**', mode: params.publish_dir_mode - publishDir "${ params.resultsDir }", + publishDir "${ publishdir }", pattern: 'Metadata/**', mode: params.publish_dir_mode label 'VV' input: + val(publishdir) path("VV_INPUT/Metadata/*") // While files from processing are staged, we instead want to use the files located in the publishDir for QC path("VV_INPUT/*") // "While files from processing are staged, we instead want to use the files located in the publishDir for QC val(skipVV) // Skips running V&V but will still publish the files @@ -49,8 +50,8 @@ process VV_AFFYMETRIX { # Export versions cat >> versions.yml < 1) { + error "Expected exactly one output directory, but found ${matches_list.size()}: ${matches_list}" + } + + def processed_dir = matches_list[0] + println "Resolved output directory: ${processed_dir}" + if (!processed_dir?.exists()) { + error "No matching output directory found (looked for ${params.accession ? 'GLDS-*' : 'results'})" + } + + // Extract GLDS accession from the directory name, if present + def dir_name = processed_dir.name + def accession_matcher = (dir_name =~ /^(GLDS-\d+)$/) + def glds_accession = accession_matcher.matches() ? accession_matcher.group(1) : null + + if (params.accession && !glds_accession) { + error "Expected a GLDS-# directory name since params.accession was set, but got: ${dir_name}" + } + + println "Resolved GLDS accession: ${glds_accession ?: 'none (results dir)'}" + + ch_processed_directory = channel.fromPath(processed_dir, checkIfExists: true) + ch_runsheet = channel.fromPath("${ processed_dir }/Metadata/*_runsheet.csv", checkIfExists: true) + ch_processing_info = channel.fromPath("$launchDir/processing_scripts/nextflow_processing_info_GLmicroarray.txt", checkIfExists: true) + ch_glds_accession = channel.value(glds_accession ?: 'NA') + + + PURGE_PROCESSING_INFO( + processed_dir, + ch_processing_info ) - GENERATE_PROTOCOL( - ch_software_versions, - ch_processing_meta, - params.skipDE + + GENERATE_MD5SUMS( + ch_processed_directory, + PURGE_PROCESSING_INFO.out.purged_processing_info ) + + def isa_file = file("${ processed_dir }/Metadata/*ISA*.zip") + if ( isa_file ) { + ch_isa = channel.fromPath("${ processed_dir }/Metadata/*ISA*.zip") + UPDATE_ASSAY_TABLE( + processed_dir, + ch_runsheet, + ch_isa, + ch_glds_accession + ) + } else { + println "${ c_back_bright_red }WARNING: No ISA archive found in ${ processed_dir }/Metadata/ -- skipping UPDATE_ASSAY_TABLE${ c_reset }" + } + } \ No newline at end of file diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/subworkflows/parse_runsheet.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/subworkflows/parse_runsheet.nf new file mode 100755 index 000000000..f459f3640 --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/subworkflows/parse_runsheet.nf @@ -0,0 +1,92 @@ +def colorCodes = [ + c_line: "┅" * 70, + c_back_bright_red: "\u001b[41;1m", + c_bright_green: "\u001b[32;1m", + c_blue: "\033[0;34m", + c_yellow: "\u001b[33;1m", + c_reset: "\033[0m" +] +// Adapted from Function: https://github.com/nf-core/rnaseq/blob/master/modules/local/process/samplesheet_check.nf +// Function to get list of [ meta, data_file ] +def get_runsheet_paths(LinkedHashMap row) { + def meta = [:] + meta.id = row["Sample Name"] + meta.biomart_id = row["biomart_attribute"] + meta.organism_sci = row.organism.replaceAll(" ","_").toLowerCase() + + // Extract factors + meta.factors = row.findAll { key, value -> + key.startsWith("Factor Value[") && key.endsWith("]") + }.collectEntries { key, value -> + [(key[13..-2]): value] // Remove "Factor Value[" and "]" + } + + meta.is_gz = row['Array Data File Path'].endsWith('.gz') + // Local staged filename: always the manufacturer's original name, minus .gz + // (decompress during staging so the QMD never has to deal with compression) + meta.file_name = meta.is_gz + ? row['Array Data File Name'].replaceAll(/\.gz$/, '') + : row['Array Data File Name'] + + return [meta, file(row['Array Data File Path'])] +} + +workflow PARSE_RUNSHEET { + take: + runsheet_path + + main: + // Process samples from the runsheet + ch_samples = runsheet_path + | splitCsv(header: true) + | map { row -> get_runsheet_paths(row) } + + ch_samples | set { ch_samples } + + // Validate consistency across samples + ch_samples + .map { meta, data_file -> [meta.biomart_id, meta.organism_sci] } + .unique() + .count() + .subscribe { count -> + if (count > 1) { + log.error "${colorCodes.c_back_bright_red}ERROR: Inconsistent metadata across samples. Please check the runsheet.${colorCodes.c_reset}" + exit 1 + } else { + println "${colorCodes.c_bright_green}Metadata consistency check passed.${colorCodes.c_reset}" + } + } + + // Print autodetected processing metadata for the first sample + ch_samples.take(1) | view { meta, data_file -> + """${colorCodes.c_bright_green}Autodetected Processing Metadata: + Biomart Attribute: ${meta.biomart_id} + Organism: ${meta.organism_sci}${colorCodes.c_reset}""" + } + + // Check that all read files are unique + ch_samples + .map { meta, data_file -> meta.file_name } + .collect() + .map { all_names -> + if (all_names.toSet().size() != all_names.size()) { + throw new RuntimeException("${colorCodes.c_back_bright_red}ERROR: Duplicate staged filenames detected — two samples would collide in the analysis work directory.${colorCodes.c_reset}") + } else { + println "${colorCodes.c_bright_green}All ${all_names.size()} staged filenames are unique.${colorCodes.c_reset}" + } + } + ch_samples + .map { meta, data_file -> data_file } + .collect() + .map { all_data_files -> + if (all_data_files.toSet().size() != all_data_files.size()) { + throw new RuntimeException("${colorCodes.c_back_bright_red}ERROR: Duplicate assay data files detected. Please check the runsheet.${colorCodes.c_reset}") + } else { + println "${colorCodes.c_bright_green}All ${all_data_files.size()} assay data files are unique.${colorCodes.c_reset}" + } + } + + emit: + samples = ch_samples + runsheet = runsheet_path +} diff --git a/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/subworkflows/stage_analysis.nf b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/subworkflows/stage_analysis.nf new file mode 100644 index 000000000..1106c91d0 --- /dev/null +++ b/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix/workflow_code/subworkflows/stage_analysis.nf @@ -0,0 +1,79 @@ +include { PARSE_RUNSHEET } from './parse_runsheet.nf' +include { FETCH_ISA } from '../modules/fetch_isa.nf' +include { ISA_TO_RUNSHEET } from '../modules/isa_to_runsheet.nf' +include { GET_ACCESSIONS } from '../modules/get_accessions.nf' +include { COPY_ARRAY_DATA_FILES } from '../modules/copy_array_data_files.nf' +//include { validateParameters; paramsSummaryLog } from 'plugin/nf-schema' +/** + * STAGE_ANALYSIS + * + * This subworkflow handles the initial setup of the Affymetrix microarray analysis: + * 1. Sets up the output directory structure + * 2. Fetches accessions if needed + * 3. Obtains or creates the runsheet + * 4. Parses the runsheet and stages array data files + */ +workflow STAGE_ANALYSIS { + take: + ch_outdir + dp_tools_plugin + accession + isa_archive_path + runsheet_path + api_url + + main: + // Parse accession, structure output directory as: + // params.outdir/ + // ├── [GLDS-#|results]/ # Main pipeline results + // └── nextflow_info/ # Pipeline execution metadata + channel.empty() | set { osd_accession } + channel.empty() | set { glds_accession } + + if ( accession ) { + GET_ACCESSIONS( accession, api_url ) + osd_accession = GET_ACCESSIONS.out.accessions_txt.map { it.readLines()[0].trim() } + glds_accession = GET_ACCESSIONS.out.accessions_txt.map { it.readLines()[1].trim() } + ch_outdir = ch_outdir.combine(glds_accession).map { outdir, glds -> "$outdir/$glds" } + } + else { + ch_outdir = ch_outdir.map { it + "/results" } + } + ch_outdir = ch_outdir.first() + + channel.empty() | set { isa_archive } + channel.empty() | set { dp_tools_version } + if ( runsheet_path == null ) { // if runsheet_path is not provided, set it up from ISA input + if ( isa_archive_path == null ) { // if isa_archive_path is not provided, fetch the ISA + FETCH_ISA( ch_outdir, osd_accession, glds_accession ) + isa_archive = FETCH_ISA.out.isa_archive + } else { + // isa_archive_path is already a channel, use it directly + isa_archive = isa_archive_path + } + ISA_TO_RUNSHEET( ch_outdir, osd_accession, glds_accession, isa_archive, dp_tools_plugin ) + runsheet_path = ISA_TO_RUNSHEET.out.runsheet + dp_tools_version = ISA_TO_RUNSHEET.out.version + } + + // Validate input parameters and runsheet + //validateParameters() + + PARSE_RUNSHEET( runsheet_path ) + samples = PARSE_RUNSHEET.out.samples + runsheet_path = PARSE_RUNSHEET.out.runsheet + + // Stage the full or truncated raw reads + COPY_ARRAY_DATA_FILES( samples ) + array_data_files = COPY_ARRAY_DATA_FILES.out + + emit: + ch_outdir = ch_outdir + samples = samples + array_data_files = array_data_files + runsheet_path = runsheet_path + isa_archive = isa_archive + osd_accession = osd_accession + glds_accession = glds_accession + dp_tools_version = dp_tools_version +} \ No newline at end of file diff --git a/Microarray/Affymetrix/Workflow_Documentation/README.md b/Microarray/Affymetrix/Workflow_Documentation/README.md index 6c28bc735..24edc6325 100644 --- a/Microarray/Affymetrix/Workflow_Documentation/README.md +++ b/Microarray/Affymetrix/Workflow_Documentation/README.md @@ -1,14 +1,14 @@ -# GeneLab RNAseq Workflow Information +# GeneLab Microarray Affymetrix (NF_MAAffymetrix) Workflow Information -> ** For the processing pipeline for Affymetrix microarray data, -[`GL-DPPD-7114.md`](../Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114.md), -GeneLab has wrapped each step of the pipeline into a workflow with validation and verification of output files built in after each step. The table below lists (and links to) each NF_MAAffymetrix version and the corresponding workflow subdirectory, the current NF_MAAffymetrix/workflow implementation is indicated. Each workflow subdirectory contains information about the workflow along with instructions for installation and usage.** +> **Starting with the processing pipeline for Affymetrix microarray data [`GL-DPPD-7114.md`](../Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114.md), +GeneLab has wrapped each step of the pipeline into a workflow with validation and verification of output files built in after each step. The table below lists (and links to) each NF_MAAffymetrix version and the corresponding workflow tag, the current NF_MAAffymetrix/workflow implementation is indicated. Each workflow tag contains information about the workflow along with instructions for installation and usage. Exact workflow run info and NF_MAAffymetrix version used to process specific datasets that have been released are available in the \*nextflow_processing_info.txt file on the [Open Science Data Repository (OSDR)](https://osdr.nasa.gov/bio/repo/), which can be found under 'Files' -> 'GeneLab Processed Microarray Data Files' -> 'Processing Info'.** ## NF_MAAffymetrix Version and Corresponding Workflow |Pipeline Version|Current Workflow Version (for respective pipeline version)|Nextflow Version| |:---------------|:---------------------------------------------------------|:---------------| -|*[GL-DPPD-7114.md](../Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114.md)|[1.0.4](NF_MAAffymetrix)|23.10.1| +|*[GL-DPPD-7114-A.md](../Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114-A.md)|[NF_MAAffymetrix_1.0.5](https://github.com/nasa/GeneLab_Data_Processing/tree/NF_MAAffymetrix_1.0.5/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix)|24.10.5| +|[GL-DPPD-7114.md](../Pipeline_GL-DPPD-7114_Versions/GL-DPPD-7114.md)|[NF_MAAffymetrix_1.0.4](https://github.com/nasa/GeneLab_Data_Processing/tree/NF_MAAffymetrix_1.0.4/Microarray/Affymetrix/Workflow_Documentation/NF_MAAffymetrix)|23.10.1| *Current GeneLab Pipeline/Workflow Implementation