# CyVerse User Guide — full corpus
Each page below begins with its canonical URL followed by its original Markdown, OKF frontmatter included. Relative links have been rewritten to absolute URLs.
---8<--- https://unm-carc.github.io/cyverse/getting-started/what-is-cyverse/
---
title: "What is CyVerse?"
description: "CyVerse's cloud cyberinfrastructure, open source software, and people: the Data Store, Discovery Environment, cloud services, funding, and how to cite it."
type: Guide
tags:
- CyVerse
- Overview
- Getting Started
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/home/what_is_cyverse.md"
title: "CyVerse Learning Materials: docs/home/what_is_cyverse.md"
author: "team:cyverse"
last_modified: "2025-03-16T09:25:29-07:00"
---
# What is CyVerse?
[CyVerse](https://cyverse.org){target=_blank} is a powerful **cloud infrastructure**, custom **open source software**, and the **_people_** who support its operations.
CyVerse is owned by The Arizona Board of Regents and is operated primarily at The University of Arizona, with partners [Texas Advanced Computing Center (TACC)](https://tacc.utexas.edu){target=_blank} at University of Texas, Austin, [Cold Spring Harbor Labs (CSHL)](https://www.cshl.edu){target=_blank} in Long Island, New York, and [Indiana University (Jetstream 2 Cloud)](https://jetstream-cloud.org){target=_blank} in Bloomington, Indiana.
CyVerse leverages public computing resources across the United States through partnerships with the National Science Foundation's [ACCESS-CI](https://access-ci.org){target=_blank} program, connected over the [Internet2](https://internet2.edu){target=_blank}
CyVerse was originally funded by the USA National Science Foundation in 2008 to handle huge datasets and complex analyses for the USA Plant Sciences community. Today, we have 16 years of experience working with tens of thousands of scientific researchers across Astronomy, Life, Earth, Health, Defense, and Space Sciences from over 160 countries.
To date, CyVerse has directly enabled over $280,000,000 in externally funded research projects in the United States, and has been cited in over 1,400 peer-reviewed research articles as the computational and data management platform which enabled the science.
Our current cyberinfrastructure platform includes:
[de]: https://unm-carc.github.io/cyverse/assets/de/logos/deIcon.svg
[data]: https://unm-carc.github.io/cyverse/assets/de/menu_items/dataIcon.svg
[atmo]: https://unm-carc.github.io/cyverse/assets/atmosphere/cacao-04.png
[ball]: https://unm-carc.github.io/cyverse/assets/de/logos/cyverse_ball_2022.png
- [![][data]{width=35}](https://unm-carc.github.io/cyverse/data-store/overview/){target=_blank} [Data Management](https://unm-carc.github.io/cyverse/data-store/overview/){target=_blank}
---
:octicons-arrow-right-24: 5GB free data storage for all basic users,
:octicons-arrow-right-24: Multi-petabyte data hosting available for [Powered By](https://unm-carc.github.io/cyverse/getting-started/powered-by/) services, sponsored projects and proposed research ([contact us](https://user.cyverse.org/requests/2){target=_blank}).
:octicons-arrow-right-24: Managed File Transfers (MFT) move your data securely with end-to-end encryption
:octicons-arrow-right-24: Store data privately, share it with your collaborators, or make it public
:octicons-arrow-right-24: Publish your data with DataCite DOI from Lyrasis
- [![][de]{width=35}](https://unm-carc.github.io/cyverse/discovery-environment/overview/){target=_blank} [Discovery Environment](https://unm-carc.github.io/cyverse/discovery-environment/overview/){target=_blank}
---
:octicons-arrow-right-24: An interactive browser-based [Data Science Workbench](https://unm-carc.github.io/cyverse/discovery-environment/using-apps/){target=_blank}
:octicons-arrow-right-24: Launch multi-core, large memory, and GPU based analyses
:octicons-arrow-right-24: Leverage major open source scientific research software applications
:octicons-arrow-right-24: Bring Your Own Software using Docker Containers
:octicons-arrow-right-24: Share your data and analyses privately and securely with collaborators
:octicons-arrow-right-24: Create and curate reproducible research objects
- [:material-cloud-tags:{ .lg .middle } Cloud Native Services](https://unm-carc.github.io/cyverse/cloud/overview/){target=_blank}
:octicons-arrow-right-24: Cloud Automation & Continuous Analysis Orchestration [CACAO](https://unm-carc.github.io/cyverse/cloud/cacao/){target=_blank} enabling cloud automation & continuous analysis orchestration on multi-cloud
:octicons-arrow-right-24: Uses templates to provision and launch scalable resources to commercial or public cloud resource providers
:octicons-arrow-right-24: Launch :simple-jupyter: Project Jupyter Hubs for tens to thousands of users in three clicks.
:octicons-arrow-right-24: Launch Open Source LLM frameworks (Ollama, OpenWebUI, DeepSeek, etc) with vector databases (Weaviate, Pinecone, etc) in secure environments and work with your data privately and securely
:octicons-arrow-right-24: [:material-gitlab: https://gitlab.com/cyverse/cacao](https://gitlab.com/cyverse/cacao){target=_blank}
[![][ball]{width=25}](https://cyverse.org/ecp){target=_blank} [Powered by](https://cyverse.org/ecp){target=_blank} - Work with our experienced data scientists and software engineers to scale your algorithms and data onto cloud and high performance compute. [Contact Us](https://user.cyverse.org/requests/3){target=_blank} if you are interested in starting an external collaborative partnership.
[![][ball]{width=25}](https://cyverse.org/teach){target=_blank} [Education and Training](https://cyverse.org/teach){target=_blank} - learn how to use containers, workflows, and public research cyberinfrastructure from our professional trainers.
[![][ball]{width=25}](https://user.cyverse.org/requests/8){target=_blank} [Cloud Resources](https://user.cyverse.org/requests/8){target=_blank} for running your class or workshop in the cloud.
-----------------------------------------------------------------------
**Funding and Citations:**
CyVerse is funded entirely by the National Science Foundation [{width="25"}](https://nsf.gov) under Award Numbers:
[{target=_blank}](https://www.nsf.gov/awardsearch/showAward?AWD_ID=0735191) [{target=_blank}](https://www.nsf.gov/awardsearch/showAward?AWD_ID=1265383) [{target=_blank}](https://www.nsf.gov/awardsearch/showAward?AWD_ID=1743442)
Please cite CyVerse appropriately when you make use of our resources, see [CyVerse citation policy](https://cyverse.org/policies/cite-cyverse){target=_blank}.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/home/what_is_cyverse.md){target=_blank} (last source update 2025-03-16), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/getting-started/account/
---
title: "Setting up your CyVerse account"
description: "Register for a CyVerse account in the User Portal, request additional services, and understand account types and subscription tiers."
type: Guide
tags:
- Accounts
- User Portal
- Getting Started
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/home/account.md"
title: "CyVerse Learning Materials: docs/home/account.md"
author: "team:cyverse"
last_modified: "2025-03-26T08:54:52-07:00"
---
# Setting up your CyVerse account
To use CyVerse's platforms, you will need to create an account.
Here's how to do so:
:octicons-arrow-right-24: Go to the [User Portal](https://user.cyverse.org/register){target=_blank} to begin the registration process.
You will be asked to enter voluntary demographic information about yourself, your contact info, and what you want to use CyVerse for.
Also, please add your ORCID to your [CyVerse User Profile](https://user.cyverse.org){target=_blank}. If you don't have an ORCID get one today!
Entering demographic information is not a requirement.
!!! warning "Avoid signing up with a personal `@gmail.com` `@hotmail.com` or private email server if possible"
CyVerse staff approve every user registration. We receive dozens of spam requests every day.
We **strongly recommended** that you use an institutional email address:`.edu`, `.org`, or `.gov` if possible. This will speed the approval process for access to certain CyVerse platforms.
Email accounts that have computer generated user names, i.e. `student1234@qq.com`, are from foreign IP addresses, or use free email services, like `qq.com`, `hotmail.com`, or `gmail.com` will be scrutinized and may be rejected.
Complete the registration process.
!!! success "Make sure to set your email address and user password before exiting the User Portal account creation wizard."
In order to reset a password, you will need to have created one in the first place.
In order to receive a password reset request, you will need to have access to the email account you signed up with.
:octicons-arrow-right-24: Check your email for an account confirmation link and follow the confirmation instructions.
Once you have confirmed your email address, you can start using your CyVerse account immediately!
!!! tip "Pop-up-blockers"
When signing up for an account, be sure that Java Script is enabled on your web browser and that any pop-up blockers are disabled.
{width="250"}
!!! tip "Check-your-spam folder"
Check your SPAM folder for the confirmation email if it does not arrive within a few minutes.
## Request Other Services
New basic account holders will have immediate access to the [Discovery Environment https://de.cyverse.org](https://unm-carc.github.io/cyverse/discovery-environment/overview/), and [Data Store](https://unm-carc.github.io/cyverse/data-store/overview/).
To register for other services and platforms, login to the [User Portal's dashboard](){target=_blank}. Under "My Services" click the 'Request Access' button next to the service(s) you would like to access. You will receive an email notification when the service is added; this may take up to 24 hours.
## Account Types
CyVerse financial sustainability model now includes a tiered subscription service where individuals can use free 'basic' tiered services for a limited amount of time.
To leverage CyVerse for research or education, you must:
(1) purchase an individual subscription (see table below),
(2) connect with us to develop an Institutional agreement (see [Professional Services](https://cyverse.org/professional-services){target=_blank} or [Powered By](https://cyverse.org/powered-by-cyverse){target=_blank}),
(3) are part of a funded research project which has an existing Professional Services agreement with CyVerse.
(4) are a current student or a faculty member at an Institution CyVerse is currently serving (specifically, the University of Arizona).
### Individual Subscriptions
In order to purchase an individual CyVerse subscription, please see [Subscribe](https://cyverse.org/subscribe){target=_blank}
**Table:** CyVerse Individual Subscription Tiers (Spring 2025)
| Service | Basic (Free) | Regular | Pro | Commercial|
|----------|--------------|---------|-----|-----------|
| Discovery Environment | Yes | Yes | Yes | Yes |
| Data Store | Yes | Yes | Yes | Yes |
| Advanced Features & APIs | - | - | Yes | Yes |
| Data Storage Limit | 5 GB | 50 GB | 4 TB | 7 TB |
| Compute Units / Year* | 200 | 1,000 | 25,000 | 250,000 |
| Access to GPU** | - | - | contact us | contact us |
| Concurrent Jobs | 1 | 2 | 4 | 8 |
| Sharing Data & Apps | None | 100 | Unlimited | Unlimited |
| Community Released Data Folder Requests | None | Yes Yes | Yes|
| DOI for Data| None | 5 | 10 | 20|
| Workshop Requests | - | - | 4 | 8 |
| Webinar Access | Yes | Yes | Yes | Yes |
| Support| Email | In-App Chat | Screen Share Support | Screen Share Support |
| Price / Year | Free | $200 | $400 | $2,400 |
!!! success "Teaching with CyVerse"
CyVerse was built as a free to use, open source cyberinfrastructure project for everyone to use. It is a privilege to offer access to the most cutting edge data science tools and computing environments in the world to students from the most under resourced and under served corners of our country with the worst internet connections.
Free "basic" account holders are intended to be undergraduates or continuing education students. The "basic" account comes with enough computing hours in the Discovery Environment for a student to complete two semester's worth (one academic year) of a courses computational assignments.
Students should be mindful of their allocation hours and use them conservatively. Analyses should not be left idle or running overnight when not in use, as they take away from the shared resource pool, and they rapidly deplete a student's free account.
Teachers should purchase a "Pro" or "Commercial" subscription, so that they can share data with their students, and request a Community Release folder, if need be.
## [Professional Services](https://cyverse.org/professional-services "Click here to learn more or request information"){target=_blank}
For over a decade CyVerse has partnered with other universities, private companies, and governmental and non-govermental organizations to provide services.
## [Powered-by CyVerse](https://cyverse.org/powered-by-cyverse "Click here to learn more about CyVerse advanced services")
Organizations may be interested in leveraging parts of CyVerse software stack or cyberinfrastructure for their own clouds, gateways, or projects.
Partnerships with CyVerse for Centers, Institutes, and Large Projects are managed through our [Powered-by](https://unm-carc.github.io/cyverse/getting-started/powered-by/) project documentation.
!!! info "Under the hood"
CyVerse accounts are managed in the User Portal and authenticated through Keycloak (with CILogon for institutional logins); see [Authentication](https://docs.cyverse.org/platform/authentication/){target=_blank} in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/home/account.md){target=_blank} (last source update 2025-03-26), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/getting-started/subscriptions/
---
title: "Individual subscriptions"
description: "CyVerse individual subscription tiers compared: storage limits, compute units, concurrent jobs, sharing, DOIs, support, and price."
type: Reference
tags:
- Subscriptions
- Accounts
- Quotas
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/home/subscriptions.md"
title: "CyVerse Learning Materials: docs/home/subscriptions.md"
author: "team:cyverse"
last_modified: "2025-03-26T08:57:29-07:00"
---
# Individual subscriptions
In order to purchase an individual CyVerse subscription, please see [Subscribe](https://cyverse.org/subscribe){target=_blank}. You can subscribe and manage billing for an existing subscription at [subscribe.cyverse.org](https://subscribe.cyverse.org){target=_blank}.
**Table:** CyVerse Individual Subscription Tiers (Spring 2025)
| Service | Basic (Free) | Regular | Pro | Commercial|
|----------|--------------|---------|-----|-----------|
| Discovery Environment | Yes | Yes | Yes | Yes |
| Data Store | Yes | Yes | Yes | Yes |
| Advanced Features & APIs | - | - | Yes | Yes |
| Data Storage Limit | 5 GB | 50 GB | 4 TB | 7 TB |
| Compute Units / Year* | 200 | 1,000 | 25,000 | 250,000 |
| Access to GPU** | - | - | contact us | contact us |
| Concurrent Jobs | 1 | 2 | 4 | 8 |
| Sharing Data & Apps | None | 100 | Unlimited | Unlimited |
| Community Released Data Folder Requests | None | Yes Yes | Yes|
| DOI for Data| None | 5 | 10 | 20|
| Workshop Requests | - | - | 4 | 8 |
| Webinar Access | Yes | Yes | Yes | Yes |
| Support| Email | In-App Chat | Screen Share Support | Screen Share Support |
| Price / Year | Free | $200 | $400 | $2,400 |
!!! success "Teaching with CyVerse"
CyVerse was built as a free to use, open source cyberinfrastructure project for everyone to use. It is a privilege to offer access to the most cutting edge data science tools and computing environments in the world to students from the most under resourced and under served corners of our country with the worst internet connections.
Free "basic" account holders are intended to be undergraduates or continuing education students. The "basic" account comes with enough computing hours in the Discovery Environment for a student to complete two semester's worth (one academic year) of a courses computational assignments.
Students should be mindful of their allocation hours and use them conservatively. Analyses should not be left idle or running overnight when not in use, as they take away from the shared resource pool, and they rapidly deplete a student's free account.
Teachers should purchase a "Pro" or "Commercial" subscription, so that they can share data with their students, and request a Community Release folder, if need be.
!!! info "Under the hood"
How subscription tiers translate into compute-hour and storage quotas is described in [Subscriptions](https://docs.cyverse.org/platform/subscriptions/){target=_blank} in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/home/subscriptions.md){target=_blank} (last source update 2025-03-26), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/getting-started/getting-help/
---
title: "Getting help"
description: "Where to find answers and how to contact CyVerse support about accounts, the Data Store, apps, and this documentation."
type: Guide
tags:
- Support
- Help
- Getting Started
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/home/getting_help.md"
title: "CyVerse Learning Materials: docs/home/getting_help.md"
author: "team:cyverse"
last_modified: "2025-03-14T07:51:00-07:00"
---
# Getting help
Most questions about CyVerse are answered in this guide or in the Discovery
Environment itself. When they are not, contact CyVerse support.
## Find an answer
* **Search this guide** with the search bar at the top of every page, or start
from the section that matches your task: [accounts](https://unm-carc.github.io/cyverse/getting-started/account/),
the [Data Store](https://unm-carc.github.io/cyverse/data-store/), the
[Discovery Environment](https://unm-carc.github.io/cyverse/discovery-environment/), or
[interactive apps](https://unm-carc.github.io/cyverse/discovery-environment/vice/).
* **Inside the Discovery Environment**, the help icon in the left sidebar opens
guides, a product tour, and contact information for the support team (see
[Logging in](https://unm-carc.github.io/cyverse/discovery-environment/login/)). For a failed or stalled
analysis, open it in the Analyses view and use its help options; see
[Managing analyses](https://unm-carc.github.io/cyverse/discovery-environment/managing-analyses/).
* **Account problems** such as password resets, service requests, and
profile updates are handled in the [CyVerse User Portal](https://user.cyverse.org){target=_blank}.
## Contact CyVerse support
* Email **** with your CyVerse username, what you were
doing, and any error message or analysis ID.
* The support channel included with your account depends on your
[subscription tier](https://unm-carc.github.io/cyverse/getting-started/subscriptions/): email for basic accounts, and in-app
chat or screen-share support for paid tiers.
* CyVerse typically performs platform maintenance on the first Tuesday of the
month, when most services are unavailable; check the date before reporting
an outage.
!!! tip "UNM CARC users"
CyVerse is one of [CARC's partner platforms](https://carc.unm.edu/docs/about/partners/){target=_blank}. For questions
about CARC clusters, including moving data between CARC storage and the
Data Store, use [CARC support](https://carc.unm.edu/docs/support/){target=_blank}.
## Issues with this documentation
This guide is maintained by UNM CARC. To report a problem or propose a change,
[open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank} or use the edit
button on any page. Questions about the original CyVerse Learning Center
material can go to .
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/home/getting_help.md){target=_blank} (last source update 2025-03-14), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/getting-started/self-guided-course/
---
title: "Self-guided course"
description: "A free CyVerse massive open online course (MOOC) covering the basics of data management and analysis on CyVerse platforms."
type: Tutorial
tags:
- Training
- Self-Guided
- Getting Started
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/home/mooc.md"
title: "CyVerse Learning Materials: docs/home/mooc.md"
author: "team:cyverse"
last_modified: "2025-03-14T07:51:00-07:00"
---
# Self-guided course
[CyVerse Self-Guided Course (USA version)](https://cyverse-learning-materials.github.io/cyverse_mooc/){target=_blank}
Created: 2022-02-14
In collaboration with CyVerse Austria and CyVerse UK, we have created a massive open online course (MOOC) for self-guided students, covering the basics of data management and analysis in CyVerse. This is a perfect starting point for new CyVerse users to get a basic understanding of how their computational workflows and data lifecycles can work in CyVerse.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/home/mooc.md){target=_blank} (last source update 2025-03-14), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/getting-started/powered-by/
---
title: "Leveraging CyVerse services"
description: "Professional Services, Powered by CyVerse, and External Collaborative Partnerships for researchers and institutions that need more than the public platform."
type: Guide
tags:
- Professional Services
- Partnerships
- CyVerse
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/home/powered_by.md"
title: "CyVerse Learning Materials: docs/home/powered_by.md"
author: "team:cyverse"
last_modified: "2025-03-14T14:47:33-07:00"
---
{ width="400" }
# Leveraging CyVerse services
CyVerse public platform provides a nominal quantity of compute and data storage to all of our users.
For researchers who need more, we also provide mechanisms for extending and powering your research.
{ width="400" }
## [Professional Services](https://cyverse.org/professional-services){target=_blank}
CyVerse offers a suite of services for institutions seeking to deploy CyVerse's cyberinfrastructure locally. With a secure, shared data management and computing environment with increased speed and performance, a local installation can better support your members' research and teaching needs for:
- Security compliance (e.g., HIPAA, ITAR, ISO, etc.)
- Federation with local and commercial cloud and high-performance computing
- Integration with local user identity management systems
Our professional services include installation, maintenance, and training for local installations of CyVerse.
Because CyVerse fully supports open source practices, you can access all of our open source infrastructure code and architecture diagrams on [Github](https://github.com/cyverse){target=_blank}.
[Contact Us](https://docs.google.com/forms/d/e/1FAIpQLSd2BqXi8DlVOeab28ZWPeUhttGqqMczMBxr8Fu1j2ud0bL3_w/viewform){target=_blank} if you're interested in knowing more.
## [Powered by CyVerse](https://cyverse.org/powered-by-cyverse){target=_blank}
Through our unique Powered by CyVerse program, third-party projects can leverage CyVerse cyberinfrastructure to provide their users with robust services such as secure single sign-on with authentication, easy and fast data transfers, access to large-scale, High-Performance Compute resources such as Jetstream2 or ACCESS-CI, as well as expertise from developers and domain scientists.
## [External Collaborative Partnership](https://cyverse.org/ecp){target=_blank}
External Collaborative Partnerships (ECP) pair you with an expert staff member to address the computational needs of your specific scientific project.
To be considered for a partnership, review the [ECP program page](https://cyverse.org/ecp){target=_blank} and complete the [ECP request form](https://user.cyverse.org/requests/3){target=_blank}. CyVerse does not provide funding support for external projects.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/home/powered_by.md){target=_blank} (last source update 2025-03-14), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/getting-started/glossary/
---
title: "Glossary & acronyms"
description: "Definitions of terms and acronyms used across CyVerse, open science, cloud computing, and research software."
type: Reference
tags:
- Glossary
- Reference
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/home/glossary.md"
title: "CyVerse Learning Materials: docs/home/glossary.md"
author: "team:cyverse"
last_modified: "2025-03-14T07:51:00-07:00"
---
# Glossary & acronyms
This glossary is to help you become more familiar with terms used in the CyVerse ecosystem as well as more broadly in Open Science and beyond.
-----------------------------------------------------------
## A
- **action:** automate a workflow in the context of CI/CD, see GitHub Actions
- **agile:** development methodology for organizing a team to complete tasks organized over short periods called 'sprints'
- **allocation:** portion of a resource assigned to a particular recipient, typical unit is a core or node hour
- **Anaconda:** open source data science platform.
- **application:** also called an 'app', a software designed to help the user to perform specific task
- **ARM:** Advanced RISC Machines, a family of central processor units
- **awesome:** a curated set of lists that provide insight into awesome software projects on GitHub
- **AVU:** Attribute-Value-Unit a components for iRODS metadata
- **AWS:** [Amazon Web Services](https://aws.amazon.com/){target=_blank} a commercial cloud
- **Azure:** [Microsoft's Azure](https://azure.microsoft.com/en-us/){target=_blank} a commercial cloud
-----------------------------------------------------------
## B
- **beta:** β, a software version which is not yet ready for publication but is being tested
- **bash:** Bash is the GNU Project's shell, the Bourne-Again Shell
- **biocontainer:** a community-driven project that provides the infrastructure and basic guidelines to create, manage and distribute bioinformatics packages (e.g conda) and containers (e.g docker, singularity)
- **bioconda:** a channel for the conda package manager specializing in bioinformatics software
-----------------------------------------------------------
## C
- **CLI:** (1) the UNIX shell [command line interface](https://en.wikipedia.org/wiki/Command-line_interface){target=_blank}, most typically Bash (2) the CyVerse Learning Institute
- **command:** a set of instructions sent to the computer, typically in a typed interface
- **conda:** an installation type of the Anaconda data science platform. Command line application for managing packages and environments
- **container:** virtualization of an operating system run within an isolated user space
- **Continuous Integration:** (CI) is testing automation to check that the application is not broken whenever new commits are integrated into the main branch
- **Continuous Delivery:** (CD) is an extension of 'continuous integration' to make sure that you can release new changes in a sustainable way
- **Continuous Deployment:** a step further than 'continuous delivery', every change that passes all stages of your production pipeline is released
- **Continuous Development:** a process for iterative software development and is an umbrella over several other processes including 'continuous integration', 'continuous testing', 'continuous delivery' and 'continuous deployment'
- **Continuous Testing:** a process of testing and automating software development
- **CPU:** central processor unit
- **CRAN:** The Comprehensive R Archive Network
- **CyVerse tool:** Software program that is integrated into the back end of the DE for use in DE apps
- **CyVerse app:** graphic interface of a tool made available for use in the DE
-----------------------------------------------------------
## D
- **Debian:** a free OS , base of other Linux distributions such as Ubuntu
- **Development:** the environment on your computer where you write code
- **DevOps** Software *Dev*elopment and information techology *Op*erations techniques for shortening the time to change software in relation to CI/CD
- **Discovery Environment (DE):** a data science workbench for running executable, interactive, and high throughput applications in the [CyVerse DE](https://de.cyverse.org){target=_blank}
- **distribution:** abbreviated as 'distro', an operating system made from a software collection based upon the Linux kernel
- **Docker:** is an open source software platform to create, deploy and manage virtualized application containers on a common operating system (OS), with an ecosystem of allied tools. A program that runs and handles life-cycle of containers and images
- **DockerHub:** an official registry of docker containers, operated by Docker.
- **DOI:** a digital object identifier. A persistant identifier number, managed by the doi.org
- **Dockerfile:** a text document that contains all the commands you would normally execute manually in order to build a Docker image. Docker can build images automatically by reading the instructions from a Dockerfile
-----------------------------------------------------------
## E
- **elastic:** (disambiguation) the ability of a cloud service provider to swiftly scale the usage of resources such as storage
- **Elastic Container Registry**
- **Elastic Container Service**
- **environment:** software that includes operating system, database system, specific tools for analysis
- **entrypoint:** In a Dockerfile, an ENTRYPOINT is an optional definition for the first part of the command to be run
-----------------------------------------------------------
## F
- **FOSS:** (1) Free and Open Source Software , (2) Foundational Open Science Skills
- **function:** a named section of a program that performs a specific task
-----------------------------------------------------------
## G
- **git:** a version control system software
- **GitHub:** a website for hosting `git` repositories -- owned by Microsoft
- **GitLab:** a website for hosting `git` repositories
- **GitOps:** using ``git`` framework as a means of deploying infrastructure on cloud using Kubernetes
- **GPU:** graphic processing unit
- **GUI:** graphical user interface
-----------------------------------------------------------
## H
- **hack:** a quick job that produces what is needed, but not well
- **HPC:** high performance computer, for large syncronous computation
- **HTC:** high throughput computer, for many parallel tasks
-----------------------------------------------------------
## I
- **IaaS:** Infrastructure as a Service . online services that provide APIs
- **iCommands:** command line application for accessing iRODS Data Store
- **IDE:** integrated development environment, typically a graphical interface for working with code language or packages
- **instance:** a single virtul machine
- **image:** self-contained, read-only ‘snapshot’ of your applications and packages, with all their dependencies
- **iRODS:** an open source integrated Rule-Oriented Data Management System
-----------------------------------------------------------
## J
- **Java:** programming language, class-based, object-oriented
- **JavaScript:** programming language
- **JSON:** Java Script Object Notation, data interchange format that uses human-readable text
- **Jupyter(Hub,Lab,Notebooks):** an IDE, originally the iPythonNotebook, operates in the browser
-----------------------------------------------------------
## K
- **kernel:** central component of most operating systems (OS)
- **Kubernetes:** an open source container orchestration platform created by Google Kubernetes is often referred to as "K8s"
-----------------------------------------------------------
## L
- **lib:** a UNIX library
- **linux:** open source Unix-like operating system
-----------------------------------------------------------
## M
- **makefile:** a file containing a set of directives used by a make build automation tool
- **markdown:** a lightweight markup language with plain text formatting syntax
- **master:** a racist word used by computer engineers to describe the `main` or `primary`` computer or branch
- **metadata::** data about data, useful for searching and querying
- **multi-thread:** a process which runs on more than one CPU or GPU core at the same time
- **Mac OS X:** Apple's popular desktop OS
-----------------------------------------------------------
## N
- **node:** a computer, typically 1 or 2 core (with many threads) server in a cloud or HPC center
-----------------------------------------------------------
## O
- **ontology:** formal naming and structural hierarchy used to describe data, also called a knowledge graph
- **organization:** a group, in the context of GitHub a place where developers contribute code to repositories
- **Operating System (OS):** software that manages computer hardware, software resources, and provides common services for computer programs
- **Open Science Grid (OSG):** national, distributed computing partnership for data-intensive research
- **ORCID:** Open Researcher and Contributor ID , a persistent digital identifier that distinguishes you from every other researcher
-----------------------------------------------------------
## P
- **PaaS:** Platform as a Service run and manage applications in cloud without complexity of developing it yourself
- **package:** an app designed for a particular langauge
- **package manager:** a collection of software tools that automates the process of installing, upgrading, configuring, and removing computer programs for a computer's operating system in a consistent manner
- **Production:** environment where users access the final code after all of the updates and testing
- **Python:** interpreted, high-level, general-purpose programming language
-----------------------------------------------------------
## Q
- **QUAY.io:** private Docker registry
-----------------------------------------------------------
## R
- **R:** data science programming language R Project
- **recipe file:** a file with installation scripts used for building software such as containers, e.g. Dockerfile
- **registry:** a storage and content delivery system, such as that used by Docker
- **remote desktop:** a VM with a graphic user interface accessed via a browser
- **repo(sitory):** a directory structure for hosting code and data
- **RST:** ReStructuredText, a markdown type file
- **ReadTheDocs:** a web service for rendering documentation (that this website uses) and readthedocs.com
- **root:** the administrative user on a linux kernel - use your powers wisely
-----------------------------------------------------------
## S
- **SaaS:** Software as a Service web based platform for using software
- **schema:** a metadata standard for labeling, tagging or coding for recording & cataloging information or structuring descriptive records. see
- **scrum:** daily set of tasks and evalautions as part of a sprint.
- **shell:** is a command line interface program that runs other programs (may be complex, technical programs or very simple programs such as making a directory). These simple, stand-alone programs are called commands
- **Singularity:** a container software, used widely on HPC, created by SyLabs
- **SLACK:** Searchable Log of All Conversation and Knowledge, a team communication tool
- **slave:** a racist word used to describe worker or secondary computers
- **sprint:** set period of time during which specific work has to be completed and made ready for review
- **Singularity def file:** (definition file) recipe for building a Singualrity container
- **Stage:** environment that is as similar to the production environment as can be for final testing
-----------------------------------------------------------
## T
- **tar:** software utility for collecting many files into one archive file, often referred to as a tarball
- **tensor:** algebraic object that describes a linear mapping from one set of algebraic objects to another
- **terminal:** a windowed emulator for directly enterinc commands to a computer
- **thread:** a CPU process or a series of linked messages in a discussion board
- **tool:** In the context of CyVerse Discovery Environment, a Docker Container
- **TPU:** tensor processing unit
- **Travis:** Travis-CI , a continuous integration software
-----------------------------------------------------------
## U
- **Ubuntu:** most popular Linux OS distribution , based on Debian
- **UNIX:** operating system
- **user:** the profile under which applications are started and run, `root` is the most powerful system administrator
-----------------------------------------------------------
## V
- **VICE:** Visual Interactive Computing Environment (see [Interactive analysis](https://unm-carc.github.io/cyverse/discovery-environment/vice/overview/))
- **virtual machine:** is a software computer that, like a physical computer, runs an operating system and applications
-----------------------------------------------------------
## W
- **waterfall:** software development broken into linear sequential phases, similar to a Gantt chart
- **W3C:** World Wide Web Consortium
- **webGL:** JavaScript API for rendering interactive 2D and 3D graphics within any compatible web browser without the use of plug-ins
- **Windows:** Microsoft's most popular desktop OS
- **workspace:** (vs. repo)
- **worker node:** A cluster typically has one or more nodes, which are the worker machines that run your containerized applications and other workloads. Each node is managed from the master, which receives updates on each node's self-reported status.
-----------------------------------------------------------
## X
- **XML:** Extensible Markup Language, data interchange format that uses human-readable text
-----------------------------------------------------------
## Y
- **YAML:** YAML Ain't Markup Language, data interchange format that uses human-readable text
-----------------------------------------------------------
## Z
- **ZenHub:** team collaboration solution built directly into GitHub that uses kanban style boards
- **Zenodo:** general-purpose open-access repository developed under the European OpenAIRE program and operated by CERN
- **zip:** a compressed file format
- **zsh:** Z-Shell now the default shell on new Mac OS X
-----------------------------------------------------------
Please cite CyVerse appropriately when you make use of our resources, see [CyVerse citation policy](https://cyverse.org/policies/cite-cyverse){target=_blank}.
**Fix or improve this documentation**
We make regular contributions to these materials, and you can suggest new materials or create and share your own.
If you have ideas or suggestions please email . You can also view, edit, and submit contributions on GitHub.
On Github: [GitHub](https://github.com/CyVerse-learning-materials){target=_blank}
Send feedback:
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/home/glossary.md){target=_blank} (last source update 2025-03-14), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/overview/
---
title: "Data Store overview"
description: "The iRODS-based CyVerse Data Store and a comparison of the ways to reach it: Discovery Environment, GoCommands, iCommands, SFTP, and WebDAV."
type: Guide
tags:
- Data Store
- iRODS
- Data Management
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/index.md"
title: "CyVerse Learning Materials: docs/ds/index.md"
author: "team:cyverse"
last_modified: "2025-04-01T14:09:40-07:00"
---
# Data Store overview
{ width="200" }
CyVerse Data Store runs the [Integrated Rule-Oriented Data System (iRODS)](https://irods.org){target=_blank} Open Source Data Management Software.
iRODS helps researchers, archivists, and others manage large, geographically dispersed computer files by providing a virtual filesystem, metadata catalog, and a rule engine to automate data management and enforce policies.
The Data Store is designed for storing, managing, and sharing your data throughout its entire lifecycle.
Integrated across all CyVerse platforms, the Data Store has high accessibility and connectivity.
The Data Store's features are aimed at helping you maintain data integrity and value, while making your data more FAIR (Findable, Accessible, Interoperable, and Reusable) with minimal effort.
This guide will walk you through the essential steps to get started, assuming you’ve already created a CyVerse account.
---
## Data Management
There are several ways to access the Data Store. These methods vary in speed, flexibility, and technical knowledge required. Different methods may suit your needs for different projects at different times.
| Method | Access Point | OS | Upload/Download | Installation/Setup Required | Account Required | Max File Size |
|------------------------|----------------------------|------------------|-----------------|-----------------------------|--------------------------|-----------------------|
| Discovery Environment | Web | Any | Both | No | Yes | 2GB/file upload, no limit for import |
| WebDAV | Web & Command line | Any | Both | No | Yes (No for public data) | No limit |
| GoCommands | Command line | Any | Both | Yes | Yes (No for public data) | No limit |
| iCommands | Command line | Linux & macOS | Both | Yes | Yes (No for public data) | No limit |
| SFTP | Desktop App & Command line | Any | Both | No (Yes for desktop app) | Yes (No for public data) | No limit |
This section covers each of the following data management methods:
1. [Discovery Environment](https://unm-carc.github.io/cyverse/data-store/de/overview/): A comprehensive web-based platform for data analysis and management
2. [GoCommands](https://unm-carc.github.io/cyverse/data-store/gocommands/overview/): A lightweight, portable command-line tool for efficient data operations on any OS
3. [iCommands](https://unm-carc.github.io/cyverse/data-store/icommands/overview/): A powerful command-line suite for advanced data management tasks on Linux
4. [SFTP](https://unm-carc.github.io/cyverse/data-store/sftp/overview/): A secure file transfer protocol accessible via command-line or GUI applications on any OS
5. [WebDAV](https://unm-carc.github.io/cyverse/data-store/webdav/overview/): A protocol extending HTTP for collaborative file management over the internet, usable on any OS
Additional resources for managing your data and team collaboration:
1. [Getting a DOI](https://unm-carc.github.io/cyverse/data-store/de/doi/): Obtain a Digital Object Identifier (DOI) for a permanent and stable link to your data
2. [Checking Data Usage](https://unm-carc.github.io/cyverse/data-store/de/check-data-usage/): Monitor your data usage and storage limits
3. [Team Access Management](https://unm-carc.github.io/cyverse/data-store/de/teams/): Create and manage teams in the Data Store for collaborative work
The [Data Store contents](https://unm-carc.github.io/cyverse/data-store/) page lists every guide in this section.
!!! info "Under the hood"
The iRODS catalog, the SFTPGo and Apache/davrods (WebDAV) services, and the HAProxy entry point at `data.cyverse.org` are described in [Data Store](https://docs.cyverse.org/platform/data-store/){target=_blank} and [Data Store administration](https://docs.cyverse.org/operations/data-store/){target=_blank} in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/index.md){target=_blank} (last source update 2025-04-01), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/de/overview/
---
title: "Data Store in the Discovery Environment"
description: "Browse, upload, inspect, delete, and search your Data Store files through the Discovery Environment web interface."
type: Guide
tags:
- Data Store
- Discovery Environment
- Data Management
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/de/index.md"
title: "CyVerse Learning Materials: docs/ds/de/index.md"
author: "team:cyverse"
last_modified: "2025-04-01T14:25:43-07:00"
---
# Data Store in the Discovery Environment
With CyVerse, you can manage data throughout the data lifecycle, from uploading, to adding metadata, to analyzing, sharing results, and making your data public for others to reuse.
The [![de]{width="25"} Discovery Environment](https://de.cyverse.org){target=_blank} interface is just one of many ways to access, view, and manage your files in the [Data Store](https://unm-carc.github.io/cyverse/data-store/overview/).
!!! question "What is the Data Store?"
The Data Store is not a separate platform; it is a service that crosscuts all of CyVerse so you can access your files (yours, shared with you, public) from anywhere in CyVerse.
[Data Store description](https://unm-carc.github.io/cyverse/data-store/overview/)
---
## :material-file-arrow-left-right-outline: Browsing Data in the Discovery Environment
[de]: https://unm-carc.github.io/cyverse/assets/de/logos/deIcon.svg
[data]: https://unm-carc.github.io/cyverse/assets/de/menu_items/dataIcon.svg
[analyses]: https://unm-carc.github.io/cyverse/assets/de/menu_items/analysisIcon.svg
[apps]: https://unm-carc.github.io/cyverse/assets/de/menu_items/appsIcon.svg
[help]: https://unm-carc.github.io/cyverse/assets/de/menu_items/helpIcon.svg
[home]: https://unm-carc.github.io/cyverse/assets/de/menu_items/homeIcon.svg
[profile]: https://unm-carc.github.io/cyverse/assets/de/icons/userIcon.svg
1. After logging in, click on the [![data]{width="25"} Data](https://de.cyverse.org/data){width="25"} icon in the left navigation menu.

To see and browse information about your data files in the Data view, press the "Customize Columns" button to select more (or fewer) columns to display, such as size, modification date, permissions, etc.
2. If the folder you're viewing has many items in it, use the < or > at the bottom of the page to move between pages. You can also change the number of items displayed per page.
3. In the upper left, you should see a "Community Data" box. If you click on this box, you can change directories to view your data files ("Home"), or data shared with your username ("Shared with me"), or files in your "Trash".
{width="200"}
As you access the folders or files within your directory, breadcrumbs near the top of the page show the folder you are viewing and its parent folder(s).
4. At the start of the breadcrumbs, you may select another root folder to view from within your home folder; click on the dropdown near your username to browse folders/files in "Shared With Me", "Community Data", or "Trash".
---
## :material-file-arrow-left-right-outline: Viewing File/Folder Details in the Discovery Environment
Both the "Details" button near the top right and the More Options menu (ellipses) at the far right in a file or folder's row allow you to view and manage several types of information about your file/folder.
You must be logged in to view file/folder details.
1. From the [![data]{width="25"} Data](https://de.cyverse.org/data){width="25"} view, click the checkbox next to a file or folder to select it and then click the "Details" or the ellipses to see specific information about the selected item, to copy the file path to the item, to add tags to the item, to edit metadata, or to set a file's info type.
2. To view your permissions on the item and those of other users, click the "Permissions" tab under "Details".
---
## :material-file-arrow-left-right-outline: Uploading Files/Folders to the Data Store via the Discovery Environment
The Data view shows a directory of the files and folders in your Data Store.
You can select an existing folder as the destination for your uploaded file(s) or click the **Folder** button to create a new folder. The default file destination is your home Data Store folder (i.e., `/iplant/home/`).
Click the "Upload" button to choose your options for importing files into the Data Store:
- To upload files from your local computer, choose **Browse Local**; a file browser will open and you may select files to upload.
- To upload files from a URL, choose **Import by URL**; you may paste in a valid HTTP or an FTP URL, then click **Import**. You may paste additional URLs or close the window by clicking **Done**.
- When you have begun the upload, you will get an automated notification that the file(s) has been queued. To view the status of an upload or import, click the **Upload** button and choose **View Upload Queue**.
---
## :material-file-arrow-left-right-outline: Deleting Files/Folders in the Discovery Environment
You must be logged in and you must own the files or folders you wish to delete.
From the [![data]{width="25"} Data](https://de.cyverse.org/data) view, select the desired file/folder by clicking the checkbox to its left. You can select multiple files/folders. To unselect an individual file/folder, click the checkbox again. You can select (or unselect) all files/folders at once by clicking the checkbox at the top of the list.
Click on the More Options menu (ellipsis) in the upper right corner of the [![data]{width="25"} Data](https://de.cyverse.org/data) view and select **Delete** from the pop-up menu. When the file has been fully deleted, you will receive an automated notification (bell icon, upper right). When deleting or moving a file/folder, you cannot change anything associated with that file/folder until you receive the completion notification.
Deleted files can be retrieved from your Trash.
!!! tip "Uploading/Importing Data via the Browser"
- You can use the DE interface to upload files of <2GB to the Data Store.
- When your Data Store file browser is open, you can also upload files from your computer by dragging them into your browser window.
- While uploading or downloading data via your browser, you must remain on the Data View until the task completes.
- The queue will only display the status of uploads from local files.
- Files imported by URL will generate an automated notification upon completion (or failure) to upload.
- When importing data from a URL, you can log out or navigate to another page or operation after you start the import; you will recieve an automated email notification when the task is complete.
- For larger files or large numbers of files, we recommend using other methods such as [SFTP](https://unm-carc.github.io/cyverse/data-store/sftp/overview/) or [GoCommands](https://unm-carc.github.io/cyverse/data-store/gocommands/overview/).
---
## :material-file-arrow-left-right-outline: Emptying Trash after Deleting Files/Folders in the Discovery Environment
Your data store allocation will not reflect deleted files/folders until you have deleted files from your Trash. To empty Trash, follow these steps:
- **Empty the Trash**: Navigate to your Trash folder. (You may find it under "Data" icon in the left navigation menu, click on the dropdown near your username to browse folders/files in "Trash")
- Once you're in the "Trash" folder, navigate to the right corner and click "Trash" followed by "Empty Trash" to permanently delete the data from your Trash folder.

- **Wait for processing**: It sometimes takes CyVerse systems a few hours to fully process file deletions, so you may not see your data storage allocation increase for several hours.
---
## :material-file-arrow-left-right-outline: Advanced Data Management Features in the Discovery Environment
The Discovery Environment also supports advanced data management tasks such as organizing your datasets, search, adding metadata to data, requesting a Digital Object Identifier (DOI), and importing or submitting data to/from NCBI SRA.
To use the Advanced Search, run a query in the Search menu, then select "Advanced Search Options".
{width="600"}
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/de/index.md){target=_blank} (last source update 2025-04-01), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/de/share/
---
title: "Sharing data in the Discovery Environment"
description: "Upload files, share them with CyVerse users or anonymously, and create public links from the Discovery Environment."
type: Guide
tags:
- Data Store
- Sharing
- Discovery Environment
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/de/share.md"
title: "CyVerse Learning Materials: docs/ds/de/share.md"
author: "team:cyverse"
last_modified: "2025-04-01T14:09:40-07:00"
---
# Sharing data in the Discovery Environment
[de]: https://unm-carc.github.io/cyverse/assets/de/logos/deIcon.svg
[data]: https://unm-carc.github.io/cyverse/assets/de/menu_items/dataIcon.svg
!!! note "This quickstart uses the Discovery Environment to upload and share data. For other ways to move data, see the [Data Store overview](https://unm-carc.github.io/cyverse/data-store/overview/)."
1. Log into the [![de]{width="25"} Discovery Environment](https://de.cyverse.org){target=_blank}.
2. Open the [![data]{width="25"} Data](https://de.cyverse.org/data){target=_blank} icon on the left.

3. Click the **Upload** button on the top right; in the dropdown menu that appears, select the preferred upload method (*Browse Local* or *Import from URL*). Additionally, you can also view your upload queue.

4. Once your file(s) is uploaded, click on the ellipses (3 dots) on the right of the file. This will open a dropdown menu with a number of options; Choose **Share**.

5. In the Share window, choose which CyVerse collaborator to share with. If your collaborator is not a registered CyVerse user, choose *anonymous*.

6. You can also generate a public URL for files, making it easier to share your files. To do so, click on the ellipses (3 dots) on the right of the file, and click Public Link(s). A window will appear with the generated URL,
which collaborators can use to download your file.
!!! warning "Generating a public URL works for files, not folders! It is suggested to compress large numbers of files prior to sharing them."

!!! info "Under the hood"
The Discovery Environment shares data through the Terrain API; the [Sharing endpoints](https://docs.cyverse.org/api/endpoints/filesystem/sharing/){target=_blank} page in the CyVerse Developer Documentation describes the calls behind the Share dialog.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/de/share.md){target=_blank} (last source update 2025-04-01), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/de/metadata/
---
title: "Adding metadata to data"
description: "View, edit, and bulk-apply metadata and metadata templates to Data Store files and folders in the Discovery Environment."
type: Guide
tags:
- Metadata
- Data Store
- Discovery Environment
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/de/metadata.md"
title: "CyVerse Learning Materials: docs/ds/de/metadata.md"
author: "team:cyverse"
last_modified: "2025-04-01T14:25:43-07:00"
---
# Adding metadata to data
CyVerse supports a variety of solutions that allow you to associate your raw data with metadata. Metadata is critically important to quality research (see this article on [FAIR Principles](https://www.nature.com/articles/sdata201618){target=_blank}), yet it is often an afterthought until you are ready to publish and share. Here are a few metadata features in CyVerse that you should know about and can adopt at the outset.
**Some things to remember about the CyVerse Discovery Environment**
- You can add metadata to a single file/folder, or in bulk to large collections of data.
- You can use your own metadata schema or apply one of several metadata templates supported in the Discovery Environment.
- Additional templates you may wish to use can be found at resources like [https://fairsharing.org/](https://fairsharing.org/){target=_blank}.
- Metadata can be managed through the DE's graphical user interface or by using iCommands at the command line.
- This guide only covers metadata options in the Discovery Environment.
---
## :material-tag-edit-outline: Viewing and Editing Metadata for a Single File/Folder in the Discovery Environment
!!! note "You must have *write* or *own* permission to edit an object's metadata."
1. Log into the [Discovery Environment](https://de.cyverse.org/de/){target=_blank}.
2. Click on the { width="25" } (Data Icon) to view or browse data. Select (checkbox) a single file/folder for which you want to add metadata.
3. Under the **More Actions** menu, click on **'Metadata'**. You will see existing metadata for the file/folder in the Attribute, Value, Unit (AVU) format.

!!! tip
A single piece of metadata, or an AVU, comprises an attribute, value, and unit. An attribute is a changeable property or characteristic of the file or folder you have selected that can be set to a value. For example, "time point" might be an attribute of a file, while '7' could be its value, and "hour" a unit of the time point.
**Adding metadata**
1. Click the "+ Add Metadata" button to add a new entry. Then follow the directions for editing metadata below.
**Editing or deleting metadata**
1. You may use the "pencil" icon to edit an existing entry or the "trash can" icon to delete an entry.
2. After you have made edits or deletions, click 'Save' to save all entries and apply the metadata.
---
## :material-tag-edit-outline: Adding Metadata to Multiple Files/Folders in the Discovery Environment
**Adding Metadata using a CyVerse Template**
1. Log into the [Discovery Environment](https://de.cyverse.org/de/){target=_blank}.
2. Click on the { width="25" } (Data Icon) to open a Data window. Select (checkbox) a single file/folder to which you want to add metadata.
3. Under the **More Actions** menu, click on **Metadata**. Click on the subsequent **More Actions** menu and select **View in Template**. You have two choices in using the template:
**A.** Choose a template; clicking **Select** will allow you to apply the template and edit the metadata manually in the DE interface.
**B.** Clicking the { width="25" }(Download icon) will download a .csv file you can edit and upload (see Applying bulk metadata below).
Click *OK* to download. (In this example, we will use the *DOI Request - DataCite Metadata*) template.
**Editing a metadata template in the DE**
Follow the steps in the "Editing or deleting metadata" from the section above.
**Editing a downloaded metadata template**
1. Unzip the downloaded template; it will contain two files: *blank.csv* and *guide.csv*. Open these files using the spreadsheet editor of your choice.
!!! tip
- *blank.csv* is the metadata template that you will complete for your data.
- *guide.csv* contains instructions for your template, and will usually include controlled vocabulary terms for metadata descriptors.
2. Edit the template in one of two ways:
1. *If all data will be in a single folder:*
- In the *blank.csv* spreadsheet, in the *'file name or path'* column, enter the file names of all the files in that folder you wish to annotate with metadata.
- In any data window, click the '⋮' (3-dots or ellipsis menu) next to any file or folder; choose 'copy path' to get the path to that item in the Data Store.
- In the remaining columns of the template, enter the values for each file/attribute combination that applies.
- If desired, add additional columns to the end of the template. The metadata in the additional columns will be saved in the Data Store but will not be stored as part of the template.
- Save the file in CSV format (i.e., [filename].csv). Avoid using spaces or specal characters when naming the file or parent folder. You may name this metadata file anything you wish, but keep it in CSV format.
2. *If data will be in multiple folders:*
- In the *blank.csv* spreadsheet, in the *'filename or path'* column, enter the full path of the top-level folder (e.g., `/iplant/home/YOURUSERNAME/FOLDERNAME`)
- In the remaining columns in the first row, enter the values for each file/attribute combination.
- Repeat for each file, making sure to add the full file path (e.g., `/iplant/home/YOURUSERNAME/FOLDERNAME`) for each file.
- If desired, add additional columns to the end of the template. The metadata in the columns will be saved in the Data Store but will not be stored as part of the template.
- Save the file in CSV format (i.e., [filename].csv). Avoid using spaces or specal characters when naming the file or parent folder. You may name this metadata file anything you wish, but keep it in CSV format.
3. In an open 'Data' window in the Discovery Environment, navigate to the appropriate location for uploading the template:
- If the first column of your metadata file contains only filenames (i.e., all data files are in the same folder), navigate to the folder and use the **Upload** button (Browse local) or your choice of upload tool to upload the metadata (csv file) to that folder.
- If the first column of your metadata file contains the full path to each file (i.e., the data files are in different folders), it does not matter where the metadata file is located on the Data Store. Use the **Upload** button (Browse local) or your choice of upload tool to upload the metadata (csv file) to an appropriate location on the Data Store.
!!! tip
For convenient management and editing, use absolute file paths
(e.g., `/iplant/home/your_file_location`) so that all of your metadata spreadsheets
will be in one location on the Data Store.
4. To apply the metadata, select (checkbox) in the Data window the name of the **folder** containing the data files to which you want to apply the metadata in bulk.
5. Click the **More Actions** menu, select 'Apply Bulk Metadata'; browse to the uploaded metadata spreadsheet and select it.
Your metadata should now be applied to your files. You should receive anotification (bell icon) in the Discovery Environment and you can confirm the
metadata have been correctly applied by following the steps in the preceding section to view metadata.
!!! info "Under the hood"
Metadata you add in the Discovery Environment is stored as iRODS attribute-value-unit (AVU) triples and managed through the Terrain API; see the [Metadata endpoints](https://docs.cyverse.org/api/endpoints/filesystem/metadata/){target=_blank} page in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/de/metadata.md){target=_blank} (last source update 2025-04-01), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/de/doi/
---
title: "Getting a DOI"
description: "Prepare, describe, and submit a dataset to CyVerse Curated Data in the Data Commons to receive a DataCite DOI."
type: Guide
tags:
- DOI
- Data Commons
- Data Publishing
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/de/doi.md"
title: "CyVerse Learning Materials: docs/ds/de/doi.md"
author: "team:cyverse"
last_modified: "2025-04-01T14:09:40-07:00"
---
# Getting a DOI
CyVerse provides Digital Object Identifiers (DOIs) for archiving research data, ensuring long-term stability and citability. DOIs are assigned through **CyVerse Curated Data in the Data Commons**, our dedicated space for preserving research outputs.
!!! note "DOIs are only available to users with a paid subscription. 'Basic' (free) CyVerse accounts do not qualify. To upgrade, explore our pricing and subscription options here: ."
## :material-ray-start-vertex-end: DOI Request Quickstart
1. ### Organize data
1. #### Create submission folder
- Organize your data so that there is one folder for each DOI (named according to the Data Commons Naming Conventions--see Step 1.b)
- Within that folder, include all files in your data package plus the ReadMe file and the inventory.
- You may have subfolders within a data package.
- You may include compressed files in a package, as described on the [FAQ](#faq), but do not compress the entire folder/package.
1. #### Name your top level folder according to the Data Commons Naming Conventions**
CyVerse Curated Data datasets are searchable and discoverable based on their metadata. While the dataset itself can have any name chosen by the creator (within reason), the folder that contains the dataset must follow the naming practices described on this page.
##### General guidelines
- Folder names must be unique.
- No invalid characters: Be sure there are **no spaces or special characters in the folder name**.
- Use underscores between each segment.
##### Folder Name Format
*\$Creator\_\$subject\_\$date*
*\$Creator*:
- The Creator entry should be the same as entered in the Creator field of the DOI request - DataCite Metadata request form.
- The Creator is the lead author, the senior author, or the organization with the primary responsibility for the dataset. Start the field (the creator's name) with a capital letter.
*Co-creators*:
- If there are two co-creators, use both names, separated by an underscore or using camel case.
- For three or more co-creators, select only one name or use a consortium name. Other contributors should be acknowledged in the metadata (as creators or contributors), which will display on the dataset landing page.
*\$subject*:
- Very briefly describes what the dataset is about.
- If the subject is more than one word, use either camel case (example: camelCase) or underscores (example: underscore_between_words) to separate the words.
- If another folder has the exact same name, you may modify the subject slightly to maintain uniqueness.
*\$date*:
- Either just the year, or the month and year, in which the dataset was created.
- Month and year should be used only if there is likely to be more than one dataset with the same creator and subject within the same year.
- Month must be a three-letter abbreviation: *Jan*, *Feb*, *Mar*, *Apr*, *May*, *Jun*, *Jul*, *Aug*, *Sep*, *Nov*, or *Dec*.
---
##### Examples
**Valid Names**
- *Walls_yam_variation_2015*
- *DeBarry_yamGenomicVariation_2016*
- *Esteva_yam_variation_Mar2016*
- *Esteva_Walls_yam_genomic_variation_Jun2016*
- *YamConsortium_Dioscorea_variation_Nov2017*
**Invalid Names**
- *WallsYamVariation_2016* (Missing underscore between the creator and the subject)
- *Esteva_yam_variation_June2016* (Month should be three letters: Jun)
- *YamConsortium_Nov2017* (No subject)
- *Walls_yam_variation_2016\#1* (Contains a special character)
- *Walls YamVariation 2020* (Contains spaces)
**Not recommended**
Although the following will pass validation, they are not recommended because the subject is too vague or too detailed:
- *Walls_variation_2016* (Subject too vague)
- *Esteva_yam_Mar2016* (Subject too vague)
- *DeBarry_yam_genetic_and_environmental_variation_with_phenotype_data_version3_Dioscorea_2016* (Too detailed)
---
1. #### Create a ReadMe file
Create a text file labeled something like *README* with the following information:
- How you obtained, organized, and labeled your dataset.
- How to reuse the data, such as which apps can analyze the data.
- The inventory (see Step 1.d) may be included as part of the ReadMe file.
- If your data include sequences, the ReadMe should include a list of corresponding BioSample IDs.
- Examples of good ReadMe files:
-
-
1. #### Create an inventory
You must create a plain text document that includes an inventory of the contents of the organized dataset (at a minimum, your dataset will contain one data file and one ReadMe file).
- The inventory may be part of the ReadMe file or a separate file.
- The inventory should include the ReadMe file and any other additional non-data materials you add to your dataset.
- If your dataset contains folders with many files (e.g., large collections of images), you do not need to list each file in the inventory. Simply describe the folder and what it contains.
- Describe the file naming conventions, if that is helpful.
---
##### Example Inventory
```text
Lyons_DOI-Example-Aug2020/: Top level directory name
README.txt: Plain text file that describes the origin of the data,
experiments, data processing, etc. Also contains a list and
description for the contents of the top-level directory
(unless a separate inventory file is provided)
License.txt: License file (e.g., GPL, MIT) that governs the use of the data
a.data1/: Directory containing data
b.data2/: Directory containing more data
c.data3/: Directory containing even more data
```
---
1. ### Add metadata
- You must provide all required metadata in the *DOI Request--Datacite 4.1* template at a minimum.
- You may add any additional metadata that is appropriate. We encourage the use of additional metadata to make your data better understood and more discoverable. For more information, including how to apply metadata, see [Adding Metadata](https://unm-carc.github.io/cyverse/data-store/de/metadata/).
!!! tip
Get recognition for your work by including [ORCIDs](https://orcid.org/){target=_blank} for yourself and all creators and contributors.
1. #### In the *Data* window, click the checkbox next to the folder.
{ width="600" }
1. #### Select *More Actions* > *Metadata*.
{ width="600" }
{ width="600" }
1. #### Select *More Actions* (again) > *View in Template*.
{ width="600" }
{ width="600" }
1. #### Choose the *DOI Request - DataCite4.1* metadata template.
{ width="600" }
{ width="600" }
1. #### Complete the required fields (marked with an asterisk) and as many of the optional fields as possible.
{ width="600" }
!!! warning
Be sure to include at least 3 subject key words or phrases, so that people can discover your data (Findability)! **Each subject should be in its own field** (click on the plus next to *Subject* to add a subject field. **DO NOT use a comma-separated list.**)
1. #### Save the template.
{ width="600" }
1. ### Submit request and wait for validations
1. #### Before you submit
Check the following to be sure everything is in order.
- There are no spaces or special characters in your file or folder names.
- You have included a ReadMe file that includes all the information specified in Step 1.c.
- You followed the Data Commons Naming Conventions
- You have filled in all the required fields in the *DOI Request - DataCite 4.1* metadata template
- You have included at least 3 subjects in your metadata
- Each subject is in a separate field (not comma-separated).
- The *description* in your metadata is adequate (other users can tell what your data describe).
- You understand that **once the DOI is issued you cannot change the data.** If you know your data will change you should consider waiting to request a DOI. If you do need to make changes later this DOI can be deprecated, a new DOI issued, and the two DOIs linked together as versions.
1. #### Submit DOI request
**In the *Data* tab, click the checkbox next to the folder.**
{ width="600" }
**Select *More Actions* > *Request DOI*.**
{ width="600" }
{ width="600" }
After verifying you have read the instructions (i.e., this guide), click *Request DOI*. You will receive a verification email that your request has been received, and a notification will be listed in the *Notifications* list in the DE.
!!! note "At this point, your folder will move to a new location under *Community Data/commons_repo/staging*."
1. #### Validations
- After submitting your request, a CyVerse curator begins validating your dataset, metadata, and overall configuration of your dataset.
- Validations are based solely on the required DOI metadata and folder-naming conventions, as well as the data's potential utility to the CyVerse and larger scientific community, not the quality of your data. This is not a peer review process.
**Possible validation actions**
- If the curator determines that minor changes are needed, they may make those changes themselves.
- If the curator determines that substantive changes are needed, they will contact you with required changes.
- If the curator determines that your dataset is not appropriate for the *Curated Data* section of the Data Commons (e.g., because it belongs in NCBI), you will be notified.
- If the curator determines that the dataset is adequately organized and the DataCite metadata are accurate, they will provide a DOI, and you will be notified of the DOI and the final dataset location.
!!! tip "To check the status of your DOI request, click *Notifications* (the bell icon) at the top right of the DE screen."
1. ### After publication
1. #### Get your dataset noticed
Metadata, the description about your data, is key to getting your dataset noticed in the world wide web. Search engines and bibliographic aggregators index the metadata that you create to obtain a DOI. Thus, it is important that you do the following:
- Make sure the metadata are complete.
- Include precise keywords in the *Subject* attribute.
- Include descriptive terms about the science and themes involved in your research. These can go in the *Subject* attribute, but you can also create additional metadata attributes specific to your dataset.
- Include methods used to generate the dataset in the *Description* attribute, and in more detail in a ReadMe file.
- Describe the dataset for a broader audience so that they understand your research. Use the *Description* field for this.
- If you or team members have an ORCID ID, make sure to include it in the metadata.
1. #### Publicize your dataset
- Consider using social media to share the DOI of your dataset, and tag CyVerse.
- If you have an interesting story about your data, contact us at , and we may be able to share it through CyVerse outreach.
- If you have a tool or workflow you developed to analyze your data in CyVerse, consider presenting it as part of our CyVerse Webinars.
---
## :material-frequently-asked-questions: FAQ
!!! question "Why should I publish my data in CyVerse Curated Data?"
CyVerse Curated Data is the ideal platform for ease of data reuse. Because it is assigned a permanent identifier (DOI), it is stable and unchangeable, making it ideal for data citation. Because the data is stored in large-scale storage resources that are monitored 24/7, it is secure. Because it allows transfer, upload, and download across different computers and platforms, it can store very large datasets. And because its data is accessible to CyVerse's suite of large-scale computational analysis resources, users can seamlessly analyze, manage, and publish new results. For more information, see [Is CyVerse Curated Data Right for My Data?](#rightformydata).
!!! question "What are the conditions for data to be published through CyVerse Curated Data?"
Several conditions must exist in your data before it can be published in CyVerse Curated Data:
- You must be a registered CyVerse account holder. To register for an account, see the [Create Account Quickstart](https://unm-carc.github.io/cyverse/getting-started/account/).
- A dataset may be up to 100 GB in size. If you are interested in depositing a larger dataset, please request an increased data allocation before requesting a permanent identifier using [this form](https://user.cyverse.org/administrative/forms/2){target=_blank}.
- Data must be both curated and static. Once the data is published, it cannot be amended (although newer versions can be published).
- Data must be organized to identify the different components (raw, preprocessed, analysis, etc.).
- Compressed files must be in [LASzip]() or open-source [gzip](http://en.wikipedia.org/wiki/Gzip){target=_blank} family of compression formats including zip, tar, or tar.gz (tgz).
- At minimum, the dataset must include a complete description according to the [DataCite standard](https://schema.datacite.org/){target=_blank}. Domain-specific schemas, however, and the addition of ReadMe files, publications, or help notes that explain the data as well as how they were obtained and can be used, are encouraged. In organizing and documenting the data, users should ask themselves, *Why would someone need to reuse this data?*
!!! question "Can I publish to the Data Commons if my data is not static and curated by CyVerse?"
Yes, you can make data available to the public via Community Released Data. You can request a community release data folder using [this form](https://user.cyverse.org/administrative/forms/7){target=_blank}.
!!! question "What is a DOI?"
A DOI is a [Digital Object Identifier](https://www.doi.org/index.html){target=_blank}. It is a permanent, redirectable identifier and URL for your dataset, so that even if the location of your dataset changes, it can still be found with the same ID. DOIs are issued by CyVerse through the [DataCite](https://datacite.org/){target=_blank} service.
!!! question "Do I need to contact CyVerse before requesting a DOI?"
The process of requesting a DOI is automated through the DE, but some tasks must be handled manually, such as DOIs for datasets with more than 1000 files or DOIs for datasets that are stored somewhere other than `/iplant/home/share/commons_repo/curated`. If you match either of those cases, please contact us at .
Also contact us if you have questions about how to organize your data or what scientific metadata to include.
!!! question "How much does a CyVerse permanent identifier cost?"
CyVerse does not charge a separate fee for each DOI; DOI requests are included with paid subscriptions (see [Individual subscriptions](https://unm-carc.github.io/cyverse/getting-started/subscriptions/)). The dataset must also meet the requirements given in the section [Is CyVerse Curated Data Right for My Data?](#rightformydata). In the future, there may be a charge for issuing permanent identifiers in the CyVerse Data Commons.
!!! question "How long will it take to obtain a permanent identifier and publish my data?"
Provided that your dataset is in good order and ready to be published, the process may take up to one week, as it may involve a dialogue with the CyVerse data curators. If your data is well organized and the metadata is complete and accurate, the process will be much faster (usually 1-2 business days). It is best to submit your request at least one week before you need the identifier (e.g., for a manuscript submission) or more for very large of complex datasets.
!!! question "Can I publish different versions of my data?"
Yes. Each new version must be documented, and will be assigned a new permanent identifier that references the original dataset. For new versions, contact us at .
!!! question "How small or big should my data be to be published?"
The size of the dataset is less important than its utility to the scientific community. Although there is no lower size limit for requesting a DOI, the default upper size limit for data allocations on CyVerse is 100 GB. If you are interested in depositing a larger dataset, please request an increased data allocation before requesting a permanent identifier.
!!! question "How do I Determine what to include?"
A data collection may be composed of multiple files and different datasets. In preparing your data for publication identify the data and other materials that you consider useful for validation and reuse of your research:
- Data associated with a research project may include multiple files with different roles.
- If there are components of your dataset that belong in a public repository such as NCBI (e.g., fastq files), submit them to the repository, rather than to CyVerse Curated Data. You may want to include a list of external files in your dataset, with links.
- Beyond data, you will include the ReadMe file (see Step 1.c), and you may include scripts or links to scripts to run your analysis. Links to analysis tools can also be included as metadata (see Step 2).
!!! question "How do I Determine how many permanent identifiers to request?"
To determine how many DOIs to request for a given data collection, consider the following:
- Size and number of components.
- How many studies or publications does it represent?
- Is your data collection formed by different datasets and are those likely to be used separately?
- Do you want to create a data collection with one DOI for the entire project and additional related DOIs for distinct datasets so that they are cited individually? DOIs can be nested, so that one dataset is part of another.
- If you are uncertain about how many DOIs to request, contact us at .
!!! question "What is the policy for submitting compressed data to CyVerse Curated Data?"
Certain file types are regularly transferred, stored, and used in applications in a compressed form, such as FASTQ for genomic data and LAZ for LIDAR data. Curated Data supports the deposition of files in the following open compressed formats: [LASzip]() and the open source [gzip]() family of compression formats including zip, tar, or tar.gz (tgz).
!!! question "Can I publish data in CyVerse if I am not a CyVerse user?"
You must have a CyVerse account to publish your data in the Data Commons repositories (Community Contributed or Curated Data). You do not have to be a user of the entire platform, but at minimum you must be able to upload data, add metadata, and use the Discovery Environment to request a DOI. If you have not used the DE's metadata features before, start with [Adding metadata to data](https://unm-carc.github.io/cyverse/data-store/de/metadata/) and read the section on metadata templates.
!!! question "How secure is the data in the Curated Data site?"
Data in our platform is stored in large-scale storage resources that are monitored 24/7. Data is authenticated through checksum analysis at ingest, and is locally and geographically replicated so that if any one system fails there will always be a safe copy of your data.
!!! question "What is CyVerse Data Commons' long-term commitment to hosting public data?"
If and when the Data Commons cannot host your data in CyVerse Curated Data, it will transfer custody of the data to another repository and will change the target URL to which the identifier points.
!!! question "What if in the future I want to move my data to another repository?"
If you want to move your data to another repository, please send a ticket with the DOI and new URL location and we will change the DOI target. You may leave a copy of the dataset in the CyVerse Curated Data site for ease of reuse within the computational environment. CyVerse will update the metadata to reflect the relationship between the two identical datasets.
!!! question "How can I make it easier for people to give me (and my co-creators) credit for using my dataset?"
Encourage others to cite your data using the DOI. Each dataset landing page includes a citation that can be copied or downloaded in standard formats (BibTEX or EndNote).
Connecting your data to your ORCID (see ) also ensures that you get credit for your work. ORCID provides a persistent digital identifier that distinguishes you from every other researcher and supports automated linkages between you and your professional activities, ensuring that your work is recognized. The DataCite metadata template includes places to list ORCIDs of the creator. The DOI creation metadata template has a place for ORCIDs of creators and contributors.
If you have published a paper that goes with your data, be sure to cite the DOI in the paper. Provide a link to the paper's DOI in the metadata under *relatedIdentifier*.
If your data include specific instructions for citing or reuse, to provide those in the ReadMe file and (if brief) in the *reuse_or_citation_conditions* metadata field.
!!! question "Who do I choose for the creator versus contributor?"
Creators are the main researchers involved in producing the data, or the authors of the publication, in priority order. To supply multiple creators, repeat this property. A creator may be a corporate/institutional or personal name; it does not need to be the person who is submitting the identifier request.
A Contributor is the institution or person(s) responsible for collecting, managing, distributing, or otherwise contributing to the development of the resource. To supply multiple contributors, repeat this property. For software, if there is an alternate entity that holds, archives, publishes, prints, distributes, releases, issues, or produces the code, use the *contributorType* *hostingInstitution* for the code repository.
You must include the role of all contributors. Choose from the dropdown list in the DOI request template.
!!! question "Which license can I use to publish my data?"
You can choose one of two open source licenses, depending on the materials you will be publishing:
- ODC PDDL for non-copyrightable materials (i.e., data only).
- CC0 for copyrightable material (Workflows, White Papers, Project Documents). If you have special circumstances that require a different license (e.g., your dataset is aggregated from previously published data that already has another license), please contact us at .
!!! question "What metadata standards does CyVerse support for data publication?"
All data will follow the DataCite metadata schema (currently using version 4.1). However, DataCite metadata is citation metadata that does not represent the complexity of the research that went behind creating your data. Therefore, we encourage you to include additional metadata. We suggest that you include the metadata records and other help documents in your publication package within a folder labeled as *metadata* so it is easily identifiable for other users. Consider taking advantage of the DE's bulk metadata application feature for adding file level metadata, especially for large datasets.
!!! question "What if I want to change or add metadata to my public data?"
If you need to make changes to the metadata of a dataset with a DOI, contact us at . If your dataset is connected to a paper that is published after the DOI is created, please contact us with the paper's DOI so we can link it in the metadata.
!!! question "Where can I go for help on permanent identifiers?"
Email the [CyVerse DOI team](mailto::doi@cyverse.org).
---
## Is the CyVerse Curated Data Repository right for my data? { #rightformydata }
**Before requesting a permanent identifier in CyVerse Curated Data through the Data Commons**, answer the following series of questions.
### Question 1. Do you have a CyVerse account?
- Are you a registered CyVerse user? If not, register at [CyVerse User Portal](https://user.cyverse.org/){target=_blank}.
- If so, have you used the Discovery Environment (DE)?
- The tools for submitting data to Data Commons Curated Data are simple to use and available as part of the DE. At a minimum, you should be able to upload and organize your data using the DE or command-line tools, and be able to apply a template.
### Question 2. Is your data ready for publication?
- Is the dataset complete, stable, and ready for public consumption?
- Are you and all contributors to the dataset prepared to move the data into the public domain (meaning that anyone can access and use the data for any purpose, including commercial purposes)? Have you sufficiently documented how the data was created such that other scientists in your field will be able to reuse it?
- If there is a standard or commonly used format for your datatype, is your data in that format? If no standard exists, is your data in a format that can be easily used by most people with open source software (e.g., tables as a CSV or text file, rather than a Microsoft Excel spreadsheet)?
- Is your data organized in a clear and reasonable structure that others will be able to understand?
If you answered *no* to any of these questions, your dataset is not yet ready for a permanent identifier through Data Commons Curated Data. Please continue to work on your dataset until it meets these requirements. Data will be reviewed by a curator to ensure that it meets these requirements.
If you would like to make your data public, but it is not complete and/or stable, you may request data hosting in the Data Commons.
### Question 3. Is your data suitable for reuse in scientific analyses?
- Is your data of the type and format that allow it to be reused in other analyses?
- Are you prepared to supply metadata for your dataset?
- Does your dataset or metadata include sufficient instructions (e.g., a ReadMe file) such that someone in your field can understand how to reuse the data?
### Question 4. Is there a canonical repository for your data?
- Does a canonical repository exist for your data? Examples include NCBI, EBI, and MG-RAST.
- If a canonical repository exists, you should use it. CyVerse is there to help fill a gap, not replace an existing resource.
### If you answered *no*
If you answered *no* to **any** of the questions above, your data may be suitable for a DOI, but not through Data Commons Curated Data. You should consider other repositories that are not geared specifically toward data analysis, such as your institution's library.
!!! tip "If your data was generated by or was input for an analysis algorithm or software that you developed yourself, please consider making the method available through CyVerse infrastructure (e.g., the Discovery Environment or Atmosphere) as well."
!!! info "Under the hood"
How curated data and DOI requests are handled is described in [Data Commons](https://docs.cyverse.org/platform/data-commons/){target=_blank} and [Permanent ID requests](https://docs.cyverse.org/api/endpoints/permanent-id-requests/){target=_blank} in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/de/doi.md){target=_blank} (last source update 2025-04-01), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/de/check-data-usage/
---
title: "Checking your data storage quotas"
description: "Check Data Store usage on the Discovery Environment dashboard and analyze folder sizes and duplicate files with the DataHog app."
type: Guide
tags:
- Data Store
- Quotas
- DataHog
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/de/check_data.md"
title: "CyVerse Learning Materials: docs/ds/de/check_data.md"
author: "team:cyverse"
last_modified: "2025-04-01T14:09:40-07:00"
---
# Checking your data storage quotas
You can see how much data you are storing in the Data Store from the Resource Usage area of the Discovery Environment Home screen. But what if you want to delete some large folders, or check for duplicate files? The DataHog app can help you understand more about the size of the folders and files you have in the Data Store.
## :material-speedometer: Resource Usage view Discovery Environment
Log into the [Discovery Environment](https://de.cyverse.org/de/){target=_blank}. After logging in, your resource usage will be displayed on the main page.
## :material-speedometer: Checking Folder Size using DataHog
1. Log into the [Discovery Environment](https://de.cyverse.org/de/){target=_blank}.
1. Click on the {width=25} to view or browse Apps.
1. In the search bar, at the top of the page, enter `DataHog`, and the click on the App suggestion that appears in the drop-down box.
1. Proceed through the app launch wizard, choosing the defaults. During the _Review and Launch_ step, push the _Launch Analysis_ button.
1. When the app is running, push the _Go to Analysis_ button.
1. Enter your CyVerse password and click _Import from iRODS_. You will also see a _CyVerse_ tab but the _iRODS_ tab is currently the preferred way to view data in CyVerse.
1. Once the import is complete, you will see a Summary of your data including a breakdown by file type, lists of files and folders by size, and lists of files and folders by date. There are other tabs to identify Duplicated Files as well as to import other storage sources so that you get a global view of your data. [Click here](https://www.youtube.com/watch?v=GQ5oMI5G9-I){target=_blank} to watch a 1-minute video on how to use DataHog.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/de/check_data.md){target=_blank} (last source update 2025-04-01), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/de/teams/
---
title: "Managing data within a team"
description: "Create public or private Teams in the Discovery Environment and share apps, tools, and data with collaborators."
type: Guide
tags:
- Teams
- Sharing
- Collaboration
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/de/teams.md"
title: "CyVerse Learning Materials: docs/ds/de/teams.md"
author: "team:cyverse"
last_modified: "2025-04-01T15:11:17-07:00"
---
# Managing data within a team
[team]: https://unm-carc.github.io/cyverse/assets/de/menu_items/teamsIcon_2.svg
The ![team]{width=20} *Teams* feature allows you to create, organize, and join public or private groups of collaborators. Teams is accessible through the left side menu, or by going through the following link: . The goal of Teams is to enable a simpler method to share apps, Tools and data with collaborators.

On the ![team]{width=20} Teams page, users are able to:
- See all public teams and teams one is part of (top left drop down menu *All Teams* :octicons-triangle-down-24:)
- Create a team (top right :fontawesome-solid-user-plus: *Team*)
---
## :material-account-group: Creating a Team
To create a team, click the :fontawesome-solid-user-plus: *Team* button on the top right. This is what the Team creation page looks like:

Here, you can:
- Pick a name of your team and add a description
- Choose whether this team is public or private
- Add (or remove) members using the *Search* bar
- Choose the user's privilege level between admin or member (Admins can allow other Members to join)
- Delete the team
!!! question "What's the difference between public and private teams?"
One can only be able to join a private team if added by the admin. Public teams allow users request to join a team through the *Join* button on the top right. Admins of that specific team will be notified.

!!! tip "Remember to save your changes before exiting the page!"
---
## :material-share-variant-outline: Sharing Apps, Tools, and Data
Being part of a team does not mean that your apps, tools and data are automatically shared. The steps below will enable sharing of apps and tools, for example, with your team, (these steps are applicable to sharing data as well):
1. Navigate to the app or tool you want to share and select it. The app/tool should be highlighted.
1. Click the :material-share-variant: *Share* button on the top of the page. This will open the *Sharing* dialog. 
1. Use the *Search* box to look for your team.
1. Select the permission level (*Read*, *Write*, or *Own*) that you want for your team and click *Done*.
Your team should now have access to the app or tool you shared!
!!! warning "When adding a new member to an existing team, everything that has already been shared with the team will be automatically be shared with the new member. *Be careful of what you share to the team!*"
---
## :material-frequently-asked-questions: FAQ
!!! question "What happens if I delete my team?"
If you delete your team, remove someone from your team or unshare any app, tool, or data, the app/tool/data will not be available for use by your collaborators.
!!! question "What about versions?"
Any changes you make to a shared app or tool will be reflected to the rest of the team. For example, if you add a version to an app, the team will be able to see the changes you have made.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/de/teams.md){target=_blank} (last source update 2025-04-01), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/gocommands/overview/
---
title: "Manage your data with GoCommands"
description: "GoCommands, a portable cross-platform command-line client for transfers, synchronization, access control, and metadata in the Data Store."
type: Guide
tags:
- GoCommands
- Data Store
- Command Line
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/gocommands/index.md"
title: "CyVerse Learning Materials: docs/ds/gocommands/index.md"
author: "team:cyverse"
last_modified: "2025-03-26T12:33:22-07:00"
---
# Manage your data with GoCommands
GoCommands is a lightweight, portable command-line tool designed for seamless data management within the Data Store. It provides comprehensive support for data transfers, bulk transfers, synchronization, access control, SFTP configuration, and metadata management. As a versatile alternative to iCommands, GoCommands runs on virtually any operating system, including embedded environments like Raspberry Pi, offering greater flexibility across platforms.
This guide covers installation, data transfer methods, access management, and metadata handling to help you efficiently manage your data in the Data Store.
The [GoCommands contents](https://unm-carc.github.io/cyverse/data-store/gocommands/) page lists every guide in this section, from installation to troubleshooting.
!!! info "Under the hood"
GoCommands talks to the Data Store's iRODS servers at `data.cyverse.org` (port 1247). The service layout is described in [Data Store](https://docs.cyverse.org/platform/data-store/){target=_blank} and [Data Store administration](https://docs.cyverse.org/operations/data-store/){target=_blank} in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/gocommands/index.md){target=_blank} (last source update 2025-03-26), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/gocommands/installation/
---
title: "GoCommands installation and upgrade"
description: "Install or upgrade GoCommands on Linux, macOS, and Windows from pre-built binaries or with conda."
type: Guide
tags:
- GoCommands
- Installation
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/gocommands/installation.md"
title: "CyVerse Learning Materials: docs/ds/gocommands/installation.md"
author: "team:cyverse"
last_modified: "2025-04-01T14:25:43-07:00"
---
# GoCommands installation and upgrade
## :material-cog-outline: Installation using Pre-built Binaries
GoCommands provides pre-built binaries for various operating systems and architectures. Choose the appropriate command for your system to install the latest version.
### :simple-apple: macOS
macOS runs on a variety of Apple devices, including MacBook, MacBook Pro, MacBook Air, iMac, Mac mini, Mac Studio, and Mac Pro. Depending on the model, it may use either an Intel/AMD 64-bit CPU or Apple Silicon (M1/M2). Follow the appropriate installation instructions based on your processor.
!!! tip "Unsure which CPU architecture your Mac uses?"
:simple-gnometerminal: Run the following command in the terminal:
```sh
uname -p
```
This will return `aarch64` / `arm64` for Apple Silicon (M1/M2) or `x86_64` for Intel-based Macs.
#### :simple-intel: Intel 64-bit
Intel processors were used in Mac devices released before 2020. If you're using an older Mac, install GoCommands with the following command:
```sh
GOCMD_VER=$(curl -L -s https://raw.githubusercontent.com/cyverse/gocommands/main/VERSION.txt); \
curl -L -s https://github.com/cyverse/gocommands/releases/download/${GOCMD_VER}/gocmd-${GOCMD_VER}-darwin-amd64.tar.gz | tar zxvf -
```
#### :simple-apple: Apple Silicon (M1/M2)
Apple introduced its custom Silicon chips (M1, M2) in 2020, replacing Intel processors in newer Mac models. If your device runs on Apple Silicon, use the following command to install GoCommands:
```sh
GOCMD_VER=$(curl -L -s https://raw.githubusercontent.com/cyverse/gocommands/main/VERSION.txt); \
curl -L -s https://github.com/cyverse/gocommands/releases/download/${GOCMD_VER}/gocmd-${GOCMD_VER}-darwin-arm64.tar.gz | tar zxvf -
```
---
### :simple-linux: Linux
Linux supports a wide range of CPU architectures. Follow the appropriate installation instructions based on your processor type.
!!! tip "Unsure which CPU architecture your system uses?"
:simple-gnometerminal: Run the following command in the terminal:
```sh
uname -p
```
This command will return the architecture type, such as `x86_64` for 64-bit Intel/AMD processors, `i386` / `i686` for 32-bit Intel/AMD processors, `aarch64` / `arm64` for 64-bit ARM processors, or `arm` for 32-bit ARM processors.
#### :simple-intel: :simple-amd: Intel/AMD 64-bit
```sh
GOCMD_VER=$(curl -L -s https://raw.githubusercontent.com/cyverse/gocommands/main/VERSION.txt); \
curl -L -s https://github.com/cyverse/gocommands/releases/download/${GOCMD_VER}/gocmd-${GOCMD_VER}-linux-amd64.tar.gz | tar zxvf -
```
#### :simple-intel: :simple-amd: Intel/AMD 32-bit
```sh
GOCMD_VER=$(curl -L -s https://raw.githubusercontent.com/cyverse/gocommands/main/VERSION.txt); \
curl -L -s https://github.com/cyverse/gocommands/releases/download/${GOCMD_VER}/gocmd-${GOCMD_VER}-linux-386.tar.gz | tar zxvf -
```
#### :simple-arm: ARM 64-bit
```sh
GOCMD_VER=$(curl -L -s https://raw.githubusercontent.com/cyverse/gocommands/main/VERSION.txt); \
curl -L -s https://github.com/cyverse/gocommands/releases/download/${GOCMD_VER}/gocmd-${GOCMD_VER}-linux-arm64.tar.gz | tar zxvf -
```
#### :simple-arm: ARM 32-bit
```sh
GOCMD_VER=$(curl -L -s https://raw.githubusercontent.com/cyverse/gocommands/main/VERSION.txt); \
curl -L -s https://github.com/cyverse/gocommands/releases/download/${GOCMD_VER}/gocmd-${GOCMD_VER}-linux-arm.tar.gz | tar zxvf -
```
---
### :material-microsoft-windows-classic: Windows
Windows primarily runs on Intel/AMD CPU architectures. Most modern systems use 64-bit Intel/AMD processors, while very old systems may run on 32-bit processors.
Windows includes two main terminal applications: `Command Prompt (CMD)` and `PowerShell`. Follow the appropriate installation instructions based on your processor type and preferred terminal.
#### :simple-gnometerminal: Command Prompt (CMD)
!!! tip "Unsure which CPU architecture your system uses?"
:simple-gnometerminal: Run the following command in the Command Prompt (CMD):
```
echo %PROCESSOR_ARCHITECTURE%
```sh
This command will return the architecture type, such as `AMD64` for 64-bit Intel/AMD processors or `x86` for 32-bit Intel/AMD processors.
##### :simple-intel: :simple-amd: Intel/AMD 64-bit
```cmd
curl -L -s -o gocmdv.txt https://raw.githubusercontent.com/cyverse/gocommands/main/VERSION.txt && set /p GOCMD_VER=Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/gocommands/installation.md){target=_blank} (last source update 2025-04-01), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/gocommands/configuration/
---
title: "GoCommands configuration"
description: "Configure GoCommands for the Data Store with the init command, existing iCommands settings, a YAML or JSON file, or environment variables."
type: Guide
tags:
- GoCommands
- Configuration
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/gocommands/configuration.md"
title: "CyVerse Learning Materials: docs/ds/gocommands/configuration.md"
author: "team:cyverse"
last_modified: "2025-04-01T14:25:43-07:00"
---
# GoCommands configuration
## :material-cog-outline: Using the `init` Command
The `init` command sets up the iRODS Host and access account for use with other GoCommands tools. Once the configuration is set, configuration files are created under the `~/.irods` directory. The configuration is fully compatible with that of iCommands.
1. **Run the following command to configure GoCommands:**
```sh
gocmd init
```
!!! tip "Getting 'Command not found error?'"
:material-braille: This error indicates that the system could not locate `gocmd` binary in the directories specified by the `$PATH` environment variable. To resolve this:
> 1. Use an absolute path: Run `./gocmd init` from the directory where you downloaded the `gocmd` binary.
> 2. For easier future use: Move the `gocmd` binary to a directory in your `$PATH`, such as `/usr/local/bin`.
> 3. Windows users: Ensure the executable is named `gocmd.exe` and run `gocmd.exe init` to initialize.
2. **Enter your Data Store account credentials when prompted. Use the following information:**
| Configuration Key | Value |
|-------------------|-------|
| `irods_host` | `data.cyverse.org` |
| `irods_port` | `1247` |
| `irods_zone_name` | `iplant` |
| `irods_user_name` | `` |
| `irods_user_password` | `` |
Use these credentials for anonymous access to the Data Store:
| Configuration Key | Value |
|-------------------|-------|
| `irods_user_name` | `anonymous` |
| `irods_user_password` | (leave empty) |
3. **To verify the current configuration, use:**
```sh
gocmd env
```
This will display the current configurations.
4. **Execute GoCommands for your task:**
```sh
gocmd ls
```
---
## :material-cog-outline: Using iCommands Configuration
GoCommands is compatible with iCommands' configuration files. It can automatically detect and use the existing iCommands configuration files located in `~/.irods`. Additionally, GoCommands creates its own configuration files in this directory, allowing users to work with both iCommands and GoCommands interchangeably.
---
## :material-cog-outline: Using an External Configuration File (YAML or JSON) without `init`
GoCommands can read configurations from YAML or JSON files without running `init` to create the `~/.irods` directory. This approach offers flexibility but requires specifying the configuration file path for each command. Here's how to use this method:
1. **Create a file named `config.yaml` using your preferred text editor:**
```sh
irods_host: "data.cyverse.org"
irods_port: 1247
irods_zone_name: "iplant"
irods_user_name: ""
irods_user_password: ""
```
!!! tip "Prefer not to include your password in the file?"
:material-security: You can omit sensitive fields like `irods_user_password`, and GoCommands will prompt you to enter the missing values during runtime.
2. **To use this configuration file, provide its path with the `-c` flag when running GoCommands:**
```sh
gocmd -c config.yaml env
```
3. **Execute GoCommands for your task:**
```sh
gocmd -c config.yaml ls
```
---
## :material-cog-outline: Using an External Configuration File (YAML or JSON)
The `init` command can be executed with an external file to automate configuration.
1. **Create a file named `config.yaml` using your preferred text editor:**
```sh
irods_host: "data.cyverse.org"
irods_port: 1247
irods_zone_name: "iplant"
irods_user_name: ""
irods_user_password: ""
```
!!! tip "Prefer not to include your password in the file?"
:material-security: You can omit sensitive fields like `irods_user_password`, and GoCommands will prompt you to enter the missing values during runtime.
2. **Execute the `init` command with the `-c` flag to configure:**
```sh
gocmd -c config.yaml init
```
---
## :material-cog-outline: Using Environmental Variables without `init`
GoCommands can read configuration directly from environmental variables, which take precedence over other configuration sources.
1. **Export the required variables in your terminal:**
```sh
export IRODS_HOST="data.cyverse.org"
export IRODS_PORT=1247
export IRODS_ZONE_NAME="iplant"
export IRODS_USER_NAME=""
export IRODS_USER_PASSWORD=""
```
!!! tip "Prefer not to set your password as an environment variable?"
:material-security: You can omit sensitive fields like `IRODS_USER_PASSWORD`, and GoCommands will prompt you to enter the missing values during runtime.
2. **Run GoCommands to verify the environment settings:**
```sh
gocmd env
```
3. **Execute GoCommands for your task:**
```sh
gocmd ls
```
---
## :material-cog-outline: Using Environmental Variables
The `init` command can be executed with environmental variables to automate configuration.
1. **Export the required variables in your terminal:**
```sh
export IRODS_HOST="data.cyverse.org"
export IRODS_PORT=1247
export IRODS_ZONE_NAME="iplant"
export IRODS_USER_NAME=""
export IRODS_USER_PASSWORD=""
```
!!! tip "Prefer not to set your password as an environment variable?"
:material-security: You can omit sensitive fields like `IRODS_USER_PASSWORD`, and GoCommands will prompt you to enter the missing values during runtime.
2. **Execute the `init` command:**
```sh
gocmd init
```
> **Note:** GoCommands will prompt you to input only the missing fields.
---
## :material-list-box-outline: Full List of Supported Configuration Fields
Below is a comprehensive list of supported fields, along with their corresponding names in JSON, YAML, and environmental variables:
| Field Name | JSON/YAML Key | Environmental Variable | Default Value |
|--------------------------------|------------------------------------|-------------------------------------|---------------------------------|
| AuthenticationScheme | `irods_authentication_scheme` | `IRODS_AUTHENTICATION_SCHEME` | native |
| AuthenticationFile | `irods_authentication_file` | `IRODS_AUTHENTICATION_FILE` | ~/irods/.irodsA |
| ClientServerNegotiation | `irods_client_server_negotiation` | `IRODS_CLIENT_SERVER_NEGOTIATION` | off |
| ClientServerPolicy | `irods_client_server_policy` | `IRODS_CLIENT_SERVER_POLICY` | CS_NEG_REFUSE |
| Host | `irods_host` | `IRODS_HOST` | |
| Port | `irods_port` | `IRODS_PORT` | 1247 |
| ZoneName | `irods_zone_name` | `IRODS_ZONE_NAME` | |
| ClientZoneName | `irods_client_zone_name` | `IRODS_CLIENT_ZONE_NAME` | |
| Username | `irods_user_name` | `IRODS_USER_NAME` | |
| ClientUsername | `irods_client_user_name` | `IRODS_CLIENT_USER_NAME` | |
| DefaultResource | `irods_default_resource` | `IRODS_DEFAULT_RESOURCE` | |
| CurrentWorkingDir | `irods_cwd` | `IRODS_CWD` | |
| Home | `irods_home` | `IRODS_HOME` | |
| DefaultHashScheme | `irods_default_hash_scheme` | `IRODS_DEFAULT_HASH_SCHEME` | SHA256 |
| MatchHashPolicy | `irods_match_hash_policy` | `IRODS_MATCH_HASH_POLICY` | |
| Debug | `irods_debug` | `IRODS_DEBUG` | |
| LogLevel | `irods_log_level` | `IRODS_LOG_LEVEL` | 0 |
| EncryptionAlgorithm | `irods_encryption_algorithm` | `IRODS_ENCRYPTION_ALGORITHM` | AES-256-CBC |
| EncryptionKeySize | `irods_encryption_key_size` | `IRODS_ENCRYPTION_KEY_SIZE` | 32 |
| EncryptionSaltSize | `irods_encryption_salt_size` | `IRODS_ENCRYPTION_SALT_SIZE` | 8 |
| EncryptionNumHashRounds | `irods_encryption_num_hash_rounds`| `IRODS_ENCRYPTION_NUM_HASH_ROUNDS` | 16 |
| SSLCACertificateFile | `irods_ssl_ca_certificate_file` | `IRODS_SSL_CA_CERTIFICATE_FILE` | |
| SSLCACertificatePath | `irods_ssl_ca_certificate_path` | `IRODS_SSL_CA_CERTIFICATE_PATH` | |
| SSLVerifyServer | `irods_ssl_verify_server` | `IRODS_SSL_VERIFY_SERVER` | hostname |
| SSLCertificateChainFile | `irods_ssl_certificate_chain_file`| `IRODS_SSL_CERTIFICATE_CHAIN_FILE` | |
| SSLCertificateKeyFile | `irods_ssl_certificate_key_file` | `IRODS_SSL_CERTIFICATE_KEY_FILE` | |
| SSLDHParamsFile | `irods_ssl_dh_params_file` | `IRODS_SSL_DH_PARAMS_FILE` | |
| Password | `irods_user_password` | `IRODS_USER_PASSWORD` | |
| Ticket | `irods_ticket` | `IRODS_TICKET` | |
| PAMToken | `irods_pam_token` | `IRODS_PAM_TOKEN` | |
| PAMTTL | `irods_pam_ttl` | `IRODS_PAM_TTL` | |
| SSLServerName | `irods_ssl_server_name` | `IRODS_SSL_SERVER_NAME` | |
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/gocommands/configuration.md){target=_blank} (last source update 2025-04-01), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/gocommands/data-management/
---
title: "Data management using GoCommands"
description: "Navigate collections and list, create, upload, download, move, rename, and remove data in the Data Store with GoCommands."
type: Guide
tags:
- GoCommands
- Data Management
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/gocommands/data_management.md"
title: "CyVerse Learning Materials: docs/ds/gocommands/data_management.md"
author: "team:cyverse"
last_modified: "2026-03-19T14:18:43-07:00"
---
# Data management using GoCommands
GoCommands offers a variety of commands to help you manage your data in the Data Store. In the Data Store, `file` and `directory` are treated as `data objects` and `collections`, respectively. It's perfectly fine to consider these terms interchangeable.
---
## :material-file-arrow-left-right-outline: Display the Current Working Collection
In the Data Store, the **current working collection** is equivalent to the concept of a current working directory in traditional file systems. You can display or change your current working collection using GoCommands.
```sh
gocmd pwd
```
By default, after configuring GoCommands, your current working collection is set to your **home directory**, which is typically located at:
```sh
//home/
```
> **Note:** Paths in the Data Store always start with the zone name `/iplant`.
---
## :material-file-arrow-left-right-outline: Change the Current Working Collection
1. **Change to a specific collection using an absolute path:**
```sh
gocmd cd /iplant/home/myUser/mydata
```
This changes your current working collection to `/iplant/home/myUser/mydata`.
2. **Use a relative path from your current location:**
Assuming your current working collection is `/iplant/home/myUser`:
```sh
gocmd cd mydata
```
3. **Return to your home collection:**
```sh
gocmd cd "~"
```
> **Note:** The `~` must be quoted to prevent shell expansion by your local shell. Without quotes, it will expand to your local machine's home directory instead of your Data Store home directory.
4. **Move up one level:**
```sh
gocmd cd ..
```
---
## :material-file-arrow-left-right-outline: List Data Objects (files) and Collections (directories)
1. **List the content of a collection:**
```sh
gocmd ls /iplant/home/myUser/mydata
```
This will display the data objects and collections in the `/iplant/home/myUser/mydata` collection:
```sh
/iplant/home/myUser/mydata:
file1.bin
file2.bin
C- /iplant/home/myUser/mydata/subdir1
```
The `C-` prefix indicates that the item is a collection (directory).
2. **List the content of the current working collection:**
```sh
gocmd ls
```
3. **List the contents of a collection in long format with additional details:**
```sh
gocmd ls -l /iplant/home/myUser/mydata
```
This command will show the data objects and collections within `/iplant/home/myUser/mydata`, along with their additional details:
```sh
/iplant/home/myUser/mydata:
myUser 0 demoRes1;rs1 436 2024-04-02.13:36 & file1.bin
myUser 1 demoRes2;rs2 436 2024-04-02.13:36 & file1.bin
myUser 0 demoRes1;rs1 700 2024-04-02.17:15 & file2.bin
myUser 1 demoRes2;rs2 700 2024-04-02.17:15 & file2.bin
C- /iplant/home/myUser/mydata/subdir1
```
- Each output line for a data object represents a replica. If iRODS is configured to create multiple replicas, you will see one line for each replica of the data object. For example, if two replicas are created, two lines will be displayed for each file.
- Each line is shown in the `owner replica_id resource_server size creation_time replica_state name` format.
- Possible replica states are:
- `&`: Good
- `X`: Stale
- `?`: Unknown
4. **List the contents of a collection with their access control lists:**
```sh
gocmd ls -A /iplant/home/myUser/mydata
```
This command will show the data objects and collections within `/iplant/home/myUser/mydata`, along with their access control lists (ACLs):
```sh
/iplant/home/myUser/mydata:
ACL - g:rodsadmin#iplant:own myUser#iplant:own
Inheritance - Disabled
file1.bin
ACL - g:rodsadmin#iplant:own myUser#iplant:own
file2.bin
ACL - g:rodsadmin#iplant:own myUser#iplant:own
C- /iplant/home/myUser/mydata/subdir1
```
- The `g:` prefix in the ACL username indicates that the user is a group.
- The ACL is displayed in the `username#zone:access_level` format.
- Most common access levels are:
- `read_object`: Allows read access to the data object or collection.
- `modify_object`: Allows modification (write) of the data object or collection.
- `own`: Grants ownership of the data object or collection.
---
## :material-file-arrow-left-right-outline: Make a Collections (directories)
1. **Create a new collection:**
```sh
gocmd mkdir /iplant/home/myUser/newCollection
```
2. **Create parent collections if they do not exist:**
```sh
gocmd mkdir -p /iplant/home/myUser/parentCollection/newCollection
```
This command creates the `newCollection` along with its parent collection `parentCollection` if it does not already exist.
---
## :material-file-arrow-left-right-outline: Upload Data Objects (files) and Collections (directories) to the Data Store
!!! warning
When uploading your data to the Data Store, avoid using:
- Spaces in names (e.g., `experiment one.fastq`)
- Special characters: \~ \`\` ! @ \# \$ % \^ & \* ( ) + = { } \[ \] \| : ; " ' < \> , ? / and \\
These may cuase issues with Discovery Environment Apps and command-line applications.
**Recommendation:** Use underscores for long names (e.g., `experiment_one.fastq`).
1. **Upload a single file:**
```sh
gocmd put /local/path/file.txt /iplant/home/myUser/
```
This command uploads the file `/local/path/file.txt` to `/iplant/home/myUser/`, creating `/iplant/home/myUser/file.txt` in the Data Store.
2. **Upload a directory and its contents:**
```sh
gocmd put /local/dir /iplant/home/myUser/
```
This command uploads the contents of the directory `/local/dir` to `/iplant/home/myUser/dir` in the Data Store. The uploaded files and subdirectories will be placed within the `/iplant/home/myUser/dir` folder.
3. **Upload with progress bars:**
```sh
gocmd put --progress /local/path/largefile.dat /iplant/home/myUser/
```
4. **Force upload:**
```sh
gocmd put -f /local/path/largefile.dat /iplant/home/myUser/
```
This command overwrites the existing file in the Data Store without prompting.
5. **Upload and verify checksum:**
```sh
gocmd put -k /local/path/important_data.txt /iplant/home/myUser/
```
This command uploads the file and verifies its integrity by calculating a checksum during transfer.
6. **Upload only different or new contents:**
```sh
gocmd put --diff /local/dir /iplant/home/myUser/
```
This command uploads only files that are different or don't exist in the destination. It compares file sizes and checksums to determine which files need updating.
7. **Upload via iCAT:**
```sh
gocmd put --icat /local/dir /iplant/home/myUser/
```
This command uses iCAT as a transfer broker. This is a default transfer method.
8. **Upload via WebDAV (HTTP):**
```sh
gocmd put --webdav /local/dir /iplant/home/myUser/
```
This command uses WebDAV (HTTP) as a transfer protocol. It is particularly useful when data transfer over port 1247 is unstable or restricted by a firewall.
---
## :material-file-arrow-left-right-outline: Download Data Objects (files) and Collections (directories) From the Data Store
1. **Download a data object to a specific local path:**
```sh
gocmd get /iplant/home/myUser/file.txt /local/path/file_new_name.txt
```
This command downloads the data object `/iplant/home/myUser/file.txt` and saves it as `/local/path/file_new_name.txt`.
2. **Download a collection to a specific local path:**
```sh
gocmd get /iplant/home/myUser/dir /local/path/
```
This command downloads the collection `/iplant/home/myUser/dir` and its contents to `/local/path`. A new directory named `dir` will be created under `/local/path`, resulting in `/local/path/dir` containing all the downloaded files and subdirectories.
3. **Download with progress bars:**
```sh
gocmd get --progress /iplant/home/myUser/largefile.dat /local/dir/
```
4. **Force download:**
```sh
gocmd get -f /iplant/home/myUser/largefile.dat .
```
This command overwrites the local file without prompting if it already exists.
5. **Download and verify checksum:**
```sh
gocmd get -k /iplant/home/myUser/important_data.txt .
```
This command downloads the file and verifies its integrity by calculating the checksum after download and comparing it with the original in the Data Store. This ensures data consistency and detects any corruption during transfer.
6. **Download only different or new contents:**
```sh
gocmd get --diff /iplant/home/myUser/dir /local/dir
```
This command downloads the source collection to the local directory, transferring only files that are different or don't exist locally. It compares file sizes and checksums to determine which files need updating, making the transfer more efficient by skipping unchanged files.
7. **Download via iCAT:**
```sh
gocmd get --icat /iplant/home/myUser/dir /local/dir
```
This command uses iCAT as a transfer broker. This is a default transfer method.
8. **Download via WebDAV (HTTP):**
```sh
gocmd get --webdav /iplant/home/myUser/dir /local/dir
```
This command uses WebDAV (HTTP) as a transfer protocol. It is particularly useful when data transfer over port 1247 is unstable or restricted by a firewall.
9. **Download with wildcard:**
```sh
gocmd get -w /iplant/home/myUser/dir/file*.txt /local/dir
```
This command downloads all data objects matching the pattern "file*.txt" from the specified collection to the local directory.
---
## :material-file-arrow-left-right-outline: Remove Data Objects (files) or Collections (directories) From the Data Store
1. **Remove a single data object:**
```sh
gocmd rm /iplant/home/myUser/file.txt
```
2. **Remove an empty collection:**
```sh
gocmd rmdir /iplant/home/myUser/emptyCollection
```
3. **Remove a collection and its contents recursively:**
```sh
gocmd rm -r /iplant/home/myUser/parentCollection
```
4. **Force remove a collection and its contents recursively:**
```sh
gocmd rm -rf /iplant/home/myUser/parentCollection
```
---
## :material-file-arrow-left-right-outline: Move/Rename Data Objects (files) or Collections (directories)
1. **Rename a data object:**
```sh
gocmd mv /iplant/home/myUser/oldfile.txt /iplant/home/myUser/newfile.txt
```
2. **Move a data object to a different collection:**
```sh
gocmd mv /iplant/home/myUser/file.txt /iplant/home/myUser/subcollection/
```
3. **Rename a collection:**
```sh
gocmd mv /iplant/home/myUser/oldcollection /iplant/home/myUser/newcollection
```
4. **Move multiple data objects with wildcard:**
```sh
gocmd mv -w /iplant/home/myUser/*.txt /iplant/home/myUser/targetcollection/
```
This command will move all data objects with the `.txt` extension from the /iplant/home/myUser/ collection to the `targetcollection`. The asterisk (*) wildcard matches any number of characters in the filename.
You can use more specific wildcard patterns for precise file selection, such as `file*.txt` to move all text files starting with "file".
---
## :material-list-box-outline: Additional Resources
For detailed information on GoCommands, refer to the [GoCommands GitHub repository](https://github.com/cyverse/gocommands/tree/main/docs/commands){target=_blank}. The following command-specific documentation is available:
- [init](https://github.com/cyverse/gocommands/tree/main/docs/commands/init.md){target=_blank}: Initialize GoCommands configuration
- [env](https://github.com/cyverse/gocommands/tree/main/docs/commands/env.md){target=_blank}: Display or modify environment variables
- [passwd](https://github.com/cyverse/gocommands/tree/main/docs/commands/passwd.md){target=_blank}: Change user password
- [cd and pwd](https://github.com/cyverse/gocommands/tree/main/docs/commands/cd_pwd.md){target=_blank}: Change and print working directory
- [ls](https://github.com/cyverse/gocommands/tree/main/docs/commands/ls.md){target=_blank}: List directory contents
- [touch](https://github.com/cyverse/gocommands/tree/main/docs/commands/touch.md){target=_blank}: Create empty files or update timestamps
- [mkdir](https://github.com/cyverse/gocommands/tree/main/docs/commands/mkdir.md){target=_blank}: Create directories
- [rm](https://github.com/cyverse/gocommands/tree/main/docs/commands/rm.md){target=_blank}: Remove files or directories
- [rmdir](https://github.com/cyverse/gocommands/tree/main/docs/commands/rmdir.md){target=_blank}: Remove directories
- [mv](https://github.com/cyverse/gocommands/tree/main/docs/commands/mv.md){target=_blank}: Move or rename files and directories
- [cp](https://github.com/cyverse/gocommands/tree/main/docs/commands/cp.md){target=_blank}: Copy files or directories
- [cat](https://github.com/cyverse/gocommands/tree/main/docs/commands/cat.md){target=_blank}: Display contents of a file
- [get](https://github.com/cyverse/gocommands/tree/main/docs/commands/get.md){target=_blank}: Download files from the Data Store
- [put](https://github.com/cyverse/gocommands/tree/main/docs/commands/put.md){target=_blank}: Upload files to the Data Store
- [bput](https://github.com/cyverse/gocommands/tree/main/docs/commands/bput.md){target=_blank}: Bulk upload files to the Data Store
- [sync](https://github.com/cyverse/gocommands/tree/main/docs/commands/sync.md){target=_blank}: Synchronize local and remote directories
- [chmod](https://github.com/cyverse/gocommands/tree/main/docs/commands/chmod.md){target=_blank}: Change access permission of files or directories
- [chmodinherit](https://github.com/cyverse/gocommands/tree/main/docs/commands/chmodinherit.md){target=_blank}: Change access permission inheritance of directories
- [lsmeta](https://github.com/cyverse/gocommands/tree/main/docs/commands/lsmeta.md){target=_blank}: List metadata of data objects, collections, resources, or users in iRODS
- [addmeta](https://github.com/cyverse/gocommands/tree/main/docs/commands/addmeta.md){target=_blank}: Add metadata to data objects, collections, resources, or users in iRODS
- [rmmeta](https://github.com/cyverse/gocommands/tree/main/docs/commands/rmmeta.md){target=_blank}: Remove metadata from data objects, collections, resources, or users in iRODS
- [copy-sftp-id](https://github.com/cyverse/gocommands/tree/main/docs/commands/copy-sftp-id.md){target=_blank}: Configure SFTP Public-key Authentication
- [svrinfo](https://github.com/cyverse/gocommands/tree/main/docs/commands/svrinfo.md){target=_blank}: Display server information
- [ps](https://github.com/cyverse/gocommands/tree/main/docs/commands/ps.md){target=_blank}: Display current iRODS sessions
- [upgrade](https://github.com/cyverse/gocommands/tree/main/docs/commands/upgrade.md){target=_blank}: Upgrade GoCommands
- [copy-sftp-id](https://github.com/cyverse/gocommands/tree/main/docs/commands/copy-sftp-id.md){target=_blank}: Copy SFTP identity file to the Data Store for SFTP public-key authentication
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/gocommands/data_management.md){target=_blank} (last source update 2026-03-19), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/gocommands/data-transfer/
---
title: "Data transfer using GoCommands"
description: "Download and upload files and folders between a local system and the Data Store with GoCommands, including thread-count tuning."
type: Guide
tags:
- GoCommands
- Data Transfer
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/gocommands/data_transfer.md"
title: "CyVerse Learning Materials: docs/ds/gocommands/data_transfer.md"
author: "team:cyverse"
last_modified: "2026-03-19T14:18:43-07:00"
---
# Data transfer using GoCommands
GoCommands provides a range of commands designed to efficiently and conveniently transfer large datasets between your local machine and the Data Store.
---
## :material-package-variant-closed: Download Data from the Data Store
1. **Download a collection to a specific local path:**
```sh
gocmd get --progress -f -k /iplant/home/myUser/mydir /local/dir/
```
This command downloads the collection `/iplant/home/myUser/mydir` and its contents to `/local/dir/`. A new directory named `mydir` will be created under `/local/dir/`, resulting in `/local/dir/mydir` containing all the downloaded files and subdirectories.
- The `--progress` flag shows progress bars, providing visual feedback during file transfers.
- The `-f` flag forces overwriting of existing local files without confirmation.
- The `-k` flag ensures file integrity by verifying checksums after the download.
2. **Sync a collection in the Data Store with a local directory:**
```sh
gocmd get --progress -f -k --diff --delete /iplant/home/myUser/mydir /local/dir/
```
This command downloads the collection `/iplant/home/myUser/mydir` and its contents to `/local/dir/`, syncing any new updates made to the original data.
- The `--diff` flag transfers only new or modified files, comparing sizes and checksums.
- The `--delete` flag removes files from the destination that don't exist in the source.
This command is equivalent to:
```sh
gocmd sync --progress -k --delete i:/iplant/home/myUser/mydir /local/dir/
```
3. **Download a collection using multiple parallel transfers:**
```sh
gocmd get --progress -f -k --thread_num 10 /iplant/home/myUser/mydir /local/dir/
```
This command downloads the collection `/iplant/home/myUser/mydir` and its contents to `/local/dir/` using 10 parallel transfer threads. Utilizing more transfer threads can maximize I/O and network bandwidth, often speeding up the transfer process significantly.
- The `--thread_num` flag sets the maximum number of threads to use for the transfer.
!!! warning "Thread Count Considerations"
1. Higher thread counts require more CPU and memory. Excessive threads may overload your system, causing performance issues. For example, Discovery Environment (DE) apps limit transfer threads to 5 due to RAM constraints.
2. The Data Store limits concurrent connections, potentially restricting high thread counts.
!!! tip "Using GoCommands on the University of Arizona Campus Network?"
When using GoCommands on the UA Campus network, include the `--webdav` flag for stable large file transfers.
---
## :material-package-variant-closed: Upload Data to the Data Store
1. **Upload a local directory to a specific path in the Data Store:**
```sh
gocmd put --progress -f -k /local/dir/ /iplant/home/myUser/mydir/
```
This command uploads the local directory `/local/dir/` and its contents to the Data Store at `/iplant/home/myUser/mydir/`. A new directory named `dir` will be created under `/iplant/home/myUser/mydir/`, resulting in `/iplant/home/myUser/mydir/dir` containing all the uploaded files and subdirectories.
- The `--progress` flag shows progress bars during file transfers.
- The `-f` flag forces overwriting of existing files in the Data Store without confirmation.
- The `-k` flag ensures file integrity by verifying checksums after upload.
2. **Sync a local directory with a collection in the Data Store:**
```sh
gocmd put --progress -f -k --diff --delete /local/dir/ /iplant/home/myUser/mydir/
```
This command uploads the local directory `/local/dir/` to `/iplant/home/myUser/mydir/` in the Data Store, syncing any new or updated files. It also removes extra files in the destination collection that do not exist locally.
- The `--diff` flag transfers only new or modified files, comparing sizes and checksums.
- The `--delete` flag removes files from the destination that don't exist in the source.
This command is equivalent to:
```sh
gocmd sync --progress -k --delete /local/dir/ i:/iplant/home/myUser/mydir/
```
3. **Upload a local directory using multiple parallel transfers:**
```sh
gocmd put --progress -f -k --thread_num 10 /local/dir/ /iplant/home/myUser/mydir/
```
This command uploads the local directory `/local/dir/` and its contents to `/iplant/home/myUser/mydir/` in the Data Store using 10 parallel transfer threads. This can significantly speed up the upload process by maximizing I/O and network bandwidth.
- The `--thread_num` flag sets the maximum number of threads to use for the transfer.
!!! warning "Thread Count Considerations"
1. Higher thread counts require more CPU and memory. Excessive threads may overload your system, causing performance issues. For example, Discovery Environment (DE) apps limit transfer threads to 5 due to RAM constraints.
2. The Data Store limits concurrent connections, potentially restricting high thread counts.
!!! tip "Using GoCommands on the University of Arizona Campus Network?"
When using GoCommands on the UA Campus network, include the `--webdav` flag for stable large file transfers.
4. **Upload a local directory containing many small files:**
```sh
gocmd bput --progress -k --thread_num 10 /local/dir/ /iplant/home/myUser/mydir/
```
This command uploads the local directory `/local/dir/` containing numerous small files to `/iplant/home/myUser/mydir/` in the Data Store. Using bundle transfer and parallel transfer can significantly speed up the transfer of many small files.
This command is equivalent to:
```sh
gocmd sync --bulk_upload --progress -k /local/dir/ i:/iplant/home/myUser/mydir/
```
!!! tip "UNM CARC users"
GoCommands is a single binary that installs without administrator rights, so
you can install it in your CARC home directory and move data directly between
CARC storage and the Data Store. CyVerse is one of
[CARC's partner platforms](https://carc.unm.edu/docs/about/partners/){target=_blank}; for moving files
onto and off CARC systems in general, see
[Transferring data](https://carc.unm.edu/docs/getting-started/transferring-data/){target=_blank} in the CARC documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/gocommands/data_transfer.md){target=_blank} (last source update 2026-03-19), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/gocommands/access-management/
---
title: "Access management using GoCommands"
description: "List and change user and group access permissions on Data Store data objects and collections with GoCommands."
type: Guide
tags:
- GoCommands
- Permissions
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/gocommands/access_management.md"
title: "CyVerse Learning Materials: docs/ds/gocommands/access_management.md"
author: "team:cyverse"
last_modified: "2025-03-25T14:01:09-07:00"
---
# Access management using GoCommands
GoCommands provides features to manage access of users and groups to data in the Data Store. The `chmod` and `chmodinherit` commands allow users to manage access permissions for data objects (files) and collections (directories). The `ls -A` command displays the current access levels assigned to users for a given file or directory.
---
## :material-security: List Access Permissions of Users for a Data Object or Collection
```sh
gocmd ls -A
```
The `-A` flag in `ls` command displays access permissions in the result.
This command will show the data objects and collections along with their access control lists (ACLs). For example:
```sh
/iplant/home/myUser/mydata:
ACL - g:rodsadmin#iplant:own myUser#iplant:own
Inheritance - Disabled
file1.bin
ACL - g:rodsadmin#iplant:own myUser#iplant:own
file2.bin
ACL - g:rodsadmin#iplant:own myUser#iplant:own
C- /iplant/home/myUser/mydata/subdir1
```
- The `g:` prefix in the ACL username indicates that the user is a group.
- The ACL is displayed in the `username#zone:access_level` format.
- Most common access levels are:
- `read_object`: Allows read access to the data object or collection.
- `modify_object`: Allows modification (write) of the data object or collection.
- `own`: Grants ownership of the data object or collection.
### Example Usage
1. **List current access levels for a data object or collection:**
```sh
gocmd ls -A /iplant/home/myUser/mydata
```
## :material-security: Change a User's or Group's Access Permission for a Data Object or Collection
```sh
gocmd chmod
```
### Access Levels
| Access Level | Description |
|-------------|-------------|
| `null` | Removes all permissions |
| `read` | Allows reading the object or collection |
| `write` | Allows reading and modifying the object or collection |
| `own` | Grants full control, including the ability to change permissions |
### Example Usage
1. **Grant a user read permission to a data object:**
```sh
gocmd chmod read anotherUser /iplant/home/myUser/file.txt
```
2. **Grant a user from a different zone read permission to a data object:**
```sh
gocmd chmod read anotherUser#anotherZone /iplant/home/myUser/file.txt
```
3. **Grant a user read permission to a collection and its contents:**
```sh
gocmd chmod -r read anotherUser /iplant/home/myUser/dir
```
4. **Grant a user write permission to a collection and its contents:**
```sh
gocmd chmod -r write anotherUser /iplant/home/myUser/dir
```
5. **Grant a user owner permission to a collection and its contents:**
```sh
gocmd chmod -r owner anotherUser /iplant/home/myUser/dir
```
6. **Remove access permission from a user to a collection and its contents:**
```sh
gocmd chmod -r none anotherUser /iplant/home/myUser/dir
```
---
## :material-security: Enable or Disable Access Permission Inheritance for a Collection
When inheritance is enabled for a collection, any new data objects or subcollections created within it will automatically inherit the same access permissions as the parent collection.
```sh
gocmd chmodinherit
```
### Inheritance Options
| Flag | Description |
|------|-------------|
| `inherit` | Enable access inheritance. Data objects and sub-collections inherit permissions from the parent collection |
| `noinherit` | Disable access inheritance. Data objects and sub-collections do not inherit permissions from the parent collection |
### Example Usage
1. **Enable inheritance for a collection:**
```sh
gocmd chmodinherit inherit /iplant/home/myUser/dir
```
2. **Disable inheritance for a collection:**
```sh
gocmd chmodinherit noinherit /iplant/home/myUser/dir
```
3. **Enable inheritance recursively for a collection and its subcollections:**
```sh
gocmd chmodinherit -r inherit /iplant/home/myUser/dir
```
4. **Disable inheritance recursively for a collection and its subcollections:**
```sh
gocmd chmodinherit -r noinherit /iplant/home/myUser/dir
```
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/gocommands/access_management.md){target=_blank} (last source update 2025-03-25), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/gocommands/metadata-management/
---
title: "Metadata management using GoCommands"
description: "List, add, and remove metadata on Data Store data objects, collections, resources, and users with GoCommands."
type: Guide
tags:
- GoCommands
- Metadata
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/gocommands/metadata_management.md"
title: "CyVerse Learning Materials: docs/ds/gocommands/metadata_management.md"
author: "team:cyverse"
last_modified: "2025-03-25T14:01:09-07:00"
---
# Metadata management using GoCommands
GoCommands provides features to manage metadata for data objects, collections, resources, and users in the Data Store using the `lsmeta`, `addmeta`, and `rmmeta` commands.
Metadata consists of three components:
- **Name (Attribute):** The name of information.
- **Value:** The actual data or information.
- **Unit (Optional):** Specifies the unit of measurement, if applicable.
---
## :material-tag-edit-outline: List Metadata of Data Objects, Collections, Resources, or Users
```sh
gocmd lsmeta [flags] ...
```
### iRODS Objects
| iROD Object | Flag | Description |
|-------------|-------------|--------|
| `data object` or `collection` | `-P` | List metadata of data objects or collections |
| `resource` | `-R` | List metadata of resources |
| `user` | `-U` | List metadata of users |
### Example Usage
1. **List metadata of a data object:**
```sh
gocmd lsmeta -P /iplant/home/myUser/file.txt
```
2. **List metadata of multiple data objects:**
```sh
gocmd lsmeta -P /iplant/home/myUser/file1.txt /iplant/home/myUser/file2.txt
```
3. **List metadata of a collection:**
```sh
gocmd lsmeta -P /iplant/home/myUser/dir
```
4. **List metadata of a resource:**
```sh
gocmd lsmeta -R myResc
```
5. **List metadata of a user:**
```sh
gocmd lsmeta -U myUser
```
## :material-tag-edit-outline: Add Metadata to Data Objects, Collections, Resources, or Users
```sh
gocmd addmeta [flags] [metadata-unit]
```
### iRODS Objects
| iROD Object | Flag | Description |
|-------------|-------------|--------|
| `data object` or `collection` | `-P` | Add metadata to a data object or collection |
| `resource` | `-R` | Add metadata to a resource |
| `user` | `-U` | Add metadata to a user |
### Example Usage
1. **Add metadata to a data object:**
```sh
gocmd addmeta -P /iplant/home/myUser/file.txt meta_name meta_value
```
1. **Add metadata to a data object with metadata-unit:**
```sh
gocmd addmeta -P /iplant/home/myUser/file.txt meta_name meta_value meta_unit
```
3. **Add metadata to a collection:**
```sh
gocmd addmeta -P /iplant/home/myUser/dir meta_name meta_value
```
4. **Add metadata to a resource:**
```sh
gocmd addmeta -R myResc meta_name meta_value
```
5. **Add metadata to a user:**
```sh
gocmd addmeta -U myUser meta_name meta_value
```
---
## :material-tag-edit-outline: Remove Metadata from Data Objects, Collections, Resources, or Users
```sh
gocmd rmmeta [flags]
```
**Note:** The `metadata-ID` is a numeric identifier for the metadata. It can be obtained from the output of the `lsmeta` command.
### iRODS Objects
| iROD Object | Flag | Description |
|-------------|-------------|--------|
| `data object` or `collection` | `-P` | Remove metadata from a data object or collection |
| `resource` | `-R` | Remove metadata from a resource |
| `user` | `-U` | Remove metadata from a user |
### Example Usage
1. **Remove metadata from a data object by name:**
```sh
gocmd rmmeta -P /iplant/home/myUser/file.txt meta_name
```
2. **Remove metadata from a data object by ID:**
```sh
gocmd rmmeta -P /iplant/home/myUser/file.txt 979206950
```
3. **Remove metadata from a collection:**
```sh
gocmd rmmeta -P /iplant/home/myUser/dir meta_name
```
4. **Remove metadata from a resource:**
```sh
gocmd rmmeta -R myResc meta_name
```
5. **Remove metadata from a user:**
```sh
gocmd rmmeta -U myUser meta_name
```
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/gocommands/metadata_management.md){target=_blank} (last source update 2025-03-25), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/gocommands/sftp-public-key-auth/
---
title: "SFTP public-key authentication configuration using GoCommands"
description: "Register an SSH public key with GoCommands for password-less SFTP access to the Data Store."
type: Guide
tags:
- GoCommands
- SFTP
- SSH Keys
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/gocommands/sftp_public_key_auth.md"
title: "CyVerse Learning Materials: docs/ds/gocommands/sftp_public_key_auth.md"
author: "team:cyverse"
last_modified: "2025-03-25T14:42:03-07:00"
---
# SFTP public-key authentication configuration using GoCommands
GoCommands provides a feature to configure public-key authentication for the Data Store's SFTP service. The `copy-sftp-id` command uploads your local SSH public keys to the Data Store, enabling password-less authentication for the SFTP service.
---
## :material-account-key-outline: Setting Up Public-key Authentication
1. **Generate an SSH key (if needed):**
```sh
ssh-keygen -t rsa -b 4096 -C "your_email@example.com"
```
This creates a private key `id_rsa` and a public key `id_rsa.pub` in the `~/.ssh` directory.
2. **To copy all SSH public keys from your `~/.ssh` directory, run:**
```sh
gocmd copy-sftp-id
```
This command automatically detects all SSH public keys for the current local user in the `~/.ssh` directory at local machine and copies them to `/iplant/home//.ssh/authorized_keys` in the Data Store. This process is similar to standard SSH public-key registration.
3. **To copy the specified SSH public key, use:**
```sh
gocmd copy-sftp-id -i ~/.ssh/id_rsa.pub
```
This command copies only the SSH public key from the `~/.ssh/id_rsa.pub` file to `/iplant/home//.ssh/authorized_keys` in the Data Store.
---
## :material-account-key-outline: Advanced Configuration
For advanced usage, you can control public-key access by manually editing the `/iplant/home//.ssh/authorized_keys` file in the Data Store. This process involves downloading the file, making changes, and then uploading it back. Here's how to do it:
1. **Download the `authorized_keys` file:**
```sh
gocmd get /iplant/home/myUser/.ssh/authorized_keys .
```
2. **Edit the file locally with a editor:**
Open it with a text editor (e.g., `vi`, `nano`):
```sh
vi authorized_keys
```
Add parameters in `key=value` format before each SSH key. Example:
```sh
expiry-time="20250320" from="10.11.12.13" ssh-rsa AAAAB3Nza... myUser
```
3. **Upload the modified file back to the Data Store:**
```sh
gocmd put authorized_keys /iplant/home/myUser/.ssh/
```
### Available Parameters
| Parameter | Description | Example |
|-------------|-------------|---------|
| `expiry-time` | Sets expiration date-time in `YYYYMMDD`, `YYYYMMDDhhmm`, or `YYYYMMDDhhmmss` format | `expiry-time="20250320"` |
| `from` | Allows access from specific IP addresses. Use IP address, CIDR, or `!` prefix to negate. Separate multiple entries with commas | `from="10.11.12.13,!10.11.12.14"` |
| `home` | Sets a specific home collection path for SFTP access within the Data Store. Use absolute path of the collection in the Data Store | `home=/iplant/home/myUser/sftp_home` |
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/gocommands/sftp_public_key_auth.md){target=_blank} (last source update 2025-03-25), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/gocommands/troubleshooting/
---
title: "GoCommands troubleshooting and issue report"
description: "Diagnose common GoCommands problems and report bugs to the GoCommands developers."
type: Guide
tags:
- GoCommands
- Troubleshooting
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/gocommands/issue_report.md"
title: "CyVerse Learning Materials: docs/ds/gocommands/issue_report.md"
author: "team:cyverse"
last_modified: "2025-03-25T14:01:09-07:00"
---
# GoCommands troubleshooting and issue report
## :material-bullseye-arrow: Troubleshooting
### Command Not Found Error
This error indicates that the system could not locate `gocmd` binary in the directories specified by the `$PATH` environment variable.
To resolve this:
1. Use an absolute or relative path: Run `./gocmd init` from the directory where you downloaded the `gocmd` binary.
2. For easier future use: Move the `gocmd` binary to a directory in your `$PATH`, such as `/usr/local/bin`.
3. Windows users: Ensure the executable is named `gocmd.exe` and run `gocmd.exe init` to initialize.
### Cannot Execute Binary File (Exec Format Error)
This error occurs when the `gocmd` binary is incompatible with your CPU architecture or operating system. To resolve this issue:
1. Review the [installation instructions](https://unm-carc.github.io/cyverse/data-store/gocommands/installation/) to ensure you downloaded the correct version for your system.
2. Verify your system's architecture and OS version:
- On Linux/macOS, use the command `uname -m` for architecture and `uname -s` for OS.
- On Windows, check System Information in the Control Panel.
3. Download the appropriate `gocmd` binary that matches your system specifications.
4. If the problem persists, seek support from the CyVerse community.
### Path Not Found Error in Windows
In Windows, the backslash (`\`) is used as the default path delimiter, while the forward slash (`/`) is used in Linux and macOS. If you encounter a "Path not found" error, ensure the following:
1. **Local Path**: Verify that your local path is correctly specified using the backslash (`\`) as the delimiter. For example:
- Correct: `C:\Users\YourName\Documents`
- Incorrect: `C:/Users/YourName/Documents`
2. **Data Store Path**: Use the forward slash (`/`) as the delimiter for paths in the Data Store. For example:
- Correct: `/iplant/home/username/folder`
- Incorrect: `\iplant\home\username\folder`
### Keep Failing Large File Transfer
If you are using GoCommands within the University of Arizona Campus or the Discovery Environment, large file transfers may fail due to campus firewall restrictions. This issue occurs when transferring files via resource servers using commands like `get`, `put`, `bput`, and `sync`. By default, these commands transfer large files (≥1GB) through resource servers or when the `--redirect` flag is specified. To avoid such failures, use the `--icat` flag to transfer data directly through ICAT.
### Request Support
If you encounter an issue that you cannot resolve, please contact [support@cyverse.org](mailto:support@cyverse.org) for assistance. Your Data Store access via GoCommands may be limited or fail due to various factors, including configuration issues, network problems, authentication errors, or data policies. The support team is available to help you identify and resolve these issues.
---
## :material-bug-check-outline: Report Bugs
If you encounter a bug, please report it to our [GitHub repository](https://github.com/cyverse/gocommands/issues){target=_blank}. Your detailed bug report is valuable in helping us improve the stability and usability of GoCommands.
When submitting a bug report, please include the following information:
- **System Information:** Specify your CPU architecture and operating system (OS).
- **Failing Command:** Provide the exact command you used that resulted in the error.
- **Debug Log:** Run the command with the `-d` flag to display debug output, then copy and paste the relevant error messages into your report. This provides valuable context for troubleshooting.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/gocommands/issue_report.md){target=_blank} (last source update 2025-03-25), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/icommands/overview/
---
title: "Manage your data with iCommands"
description: "iCommands, the iRODS Consortium's command-line suite for transfers, synchronization, access control, and metadata in the Data Store."
type: Guide
tags:
- iCommands
- iRODS
- Data Store
- Command Line
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/icommands/index.md"
title: "CyVerse Learning Materials: docs/ds/icommands/index.md"
author: "team:cyverse"
last_modified: "2025-03-20T16:45:42-07:00"
---
# Manage your data with iCommands
iCommands is a command-line tool designed for efficient data management in iRODS. Developed by the iRODS Consortium, which created the iRODS data system powering the Data Store, iCommands provides robust support for data transfers, synchronization, access control, and metadata management. It is compatible with popular Linux environments.
This guide covers installation, data transfer, access management, and metadata handling to help you efficiently manage your data in the Data Store.
The [iCommands contents](https://unm-carc.github.io/cyverse/data-store/icommands/) page lists every guide in this section, from installation to troubleshooting.
!!! info "Under the hood"
iCommands talks to the Data Store's iRODS servers at `data.cyverse.org` (port 1247). The service layout is described in [Data Store](https://docs.cyverse.org/platform/data-store/){target=_blank} and [Data Store administration](https://docs.cyverse.org/operations/data-store/){target=_blank} in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/icommands/index.md){target=_blank} (last source update 2025-03-20), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/icommands/installation/
---
title: "iCommands installation and upgrade"
description: "Install iCommands on Linux with APT, YUM, or Zypper, or manually for a specific version."
type: Guide
tags:
- iCommands
- Installation
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/icommands/installation.md"
title: "CyVerse Learning Materials: docs/ds/icommands/installation.md"
author: "team:cyverse"
last_modified: "2025-03-27T12:34:29-07:00"
---
# iCommands installation and upgrade
## :material-cog-outline: Installing iCommands Using Linux Package Managers
iCommands can be installed using popular Linux package managers such as `apt`, `yum`, or `zypper`. Follow the instructions below to install the latest version of iCommands based on your system's package manager.
> **Note:** You need administrative privileges to install the iCommands package on the system.
### :simple-ubuntu: APT (Debian/Ubuntu)
APT is the default package manager for Debian-based distributions like Ubuntu. Use the following commands to add the iRODS repository, import the signing key, and update your package list:
```sh
wget -qO - https://packages.irods.org/irods-signing-key.asc | sudo apt-key add -
echo "deb [arch=amd64] https://packages.irods.org/apt/ $(lsb_release -sc) main" | sudo tee /etc/apt/sources.list.d/renci-irods.list
sudo apt-get update
```
Then use the following command to install iCommands:
```sh
sudo apt install irods-icommands
```
---
### :simple-redhat: YUM (RHEL/CentOS/Fedora)
YUM is the default package manager for Red Hat-based distributions such as RHEL, CentOS, and Fedora. To install the public signing key and add the iRODS repository, execute the following commands:
```sh
sudo rpm --import https://packages.irods.org/irods-signing-key.asc
wget -qO - https://packages.irods.org/renci-irods.yum.repo | sudo tee /etc/yum.repos.d/renci-irods.yum.repo
```
Then use the following command to install iCommands:
```sh
sudo yum install irods-icommands
```
---
### :simple-suse: ZYPPER (openSUSE/SUSE Linux Enterprise)
ZYPPER is the package manager for SUSE-based distributions like openSUSE and SUSE Linux Enterprise. Use these commands to import the signing key and add the iRODS repository:
```sh
sudo rpm --import https://packages.irods.org/irods-signing-key.asc
wget -qO - https://packages.irods.org/renci-irods.zypp.repo | sudo tee /etc/zypp/repos.d/renci-irods.zypp.repo
```
Then use the following command to install iCommands:
```sh
sudo zypper install irods-icommands
```
---
## :material-braille: Manual Installation or Specific Versions
If you prefer to manually download binaries or require a specific version of iCommands, visit the official iRODS download page for more information:
[iRODS Download Page](https://irods.org/download/){target=_blank}
Alternatively, you can directly browse the repositories for specific versions:
- **APT Repository**: [https://packages.irods.org/apt/pool/](https://packages.irods.org/apt/pool/){target=_blank}
- **YUM/ZYPPER Repository**: [https://packages.irods.org/yum/pool](https://packages.irods.org/yum/pool){target=_blank}
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/icommands/installation.md){target=_blank} (last source update 2025-03-27), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/icommands/configuration/
---
title: "iCommands configuration"
description: "Configure iCommands for the Data Store with iinit or environment variables."
type: Guide
tags:
- iCommands
- Configuration
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/icommands/configuration.md"
title: "CyVerse Learning Materials: docs/ds/icommands/configuration.md"
author: "team:cyverse"
last_modified: "2025-04-01T14:25:43-07:00"
---
# iCommands configuration
## :material-cog-outline: Using the `iinit` Command
The `iinit` command sets up the iRODS Host and access account for use with other iCommands tools. Once the configuration is set, configuration files are created under the `~/.irods` directory.
!!! tip "Running iCommands in an HPC environment?"
:material-braille: To use pre-installed iCommands in an HPC environment, run following to access iCommands:
```sh
module load irods
```
1. **Run the following command to configure iCommands:**
```sh
iinit
```
2. **Enter your Data Store account credentials when prompted. Use the following information:**
| Configuration Key | Value |
|-------------------|-------|
| `irods_host` | `data.cyverse.org` |
| `irods_port` | `1247` |
| `irods_zone_name` | `iplant` |
| `irods_user_name` | `` |
| `irods_user_password` | `` |
Use these credentials for anonymous access to the Data Store:
| Configuration Key | Value |
|-------------------|-------|
| `irods_user_name` | `anonymous` |
| `irods_user_password` | (leave empty) |
3. **To verify the current configuration, use:**
```sh
ienv
```
This will display the current configurations.
4. **Execute iCommands for your task:**
```sh
ils
```
---
## :material-cog-outline: Using Environmental Variables
The `iinit` command can be executed with environmental variables to automate configuration.
1. **Export the required variables in your terminal:**
```sh
export IRODS_HOST="data.cyverse.org"
export IRODS_PORT=1247
export IRODS_ZONE_NAME="iplant"
export IRODS_USER_NAME=""
```
> **Note:** iCommands does not support setting passwords via environment variables.
2. **Execute the `iinit` command:**
```sh
iinit
```
> **Note:** `iinit` will prompt you to input only the missing fields.
---
## :material-list-box-outline: Full List of Supported Configuration Fields
Below is a comprehensive list of supported fields, along with their corresponding names in JSON, and environmental variables:
| Field Name | JSON Key | Environmental Variable | Default Value |
|--------------------------------|------------------------------------|-------------------------------------|---------------------------------|
| AuthenticationScheme | `irods_authentication_scheme` | `IRODS_AUTHENTICATION_SCHEME` | native |
| AuthenticationFile | `irods_authentication_file` | `IRODS_AUTHENTICATION_FILE` | ~/irods/.irodsA |
| ClientServerNegotiation | `irods_client_server_negotiation` | `IRODS_CLIENT_SERVER_NEGOTIATION` | off |
| ClientServerPolicy | `irods_client_server_policy` | `IRODS_CLIENT_SERVER_POLICY` | CS_NEG_REFUSE |
| Host | `irods_host` | `IRODS_HOST` | |
| Port | `irods_port` | `IRODS_PORT` | 1247 |
| ZoneName | `irods_zone_name` | `IRODS_ZONE_NAME` | |
| ClientZoneName | `irods_client_zone_name` | `IRODS_CLIENT_ZONE_NAME` | |
| Username | `irods_user_name` | `IRODS_USER_NAME` | |
| ClientUsername | `irods_client_user_name` | `IRODS_CLIENT_USER_NAME` | |
| DefaultResource | `irods_default_resource` | `IRODS_DEFAULT_RESOURCE` | |
| CurrentWorkingDir | `irods_cwd` | `IRODS_CWD` | |
| Home | `irods_home` | `IRODS_HOME` | |
| DefaultHashScheme | `irods_default_hash_scheme` | `IRODS_DEFAULT_HASH_SCHEME` | SHA256 |
| MatchHashPolicy | `irods_match_hash_policy` | `IRODS_MATCH_HASH_POLICY` | |
| Debug | `irods_debug` | `IRODS_DEBUG` | |
| LogLevel | `irods_log_level` | `IRODS_LOG_LEVEL` | 0 |
| EncryptionAlgorithm | `irods_encryption_algorithm` | `IRODS_ENCRYPTION_ALGORITHM` | AES-256-CBC |
| EncryptionKeySize | `irods_encryption_key_size` | `IRODS_ENCRYPTION_KEY_SIZE` | 32 |
| EncryptionSaltSize | `irods_encryption_salt_size` | `IRODS_ENCRYPTION_SALT_SIZE` | 8 |
| EncryptionNumHashRounds | `irods_encryption_num_hash_rounds`| `IRODS_ENCRYPTION_NUM_HASH_ROUNDS` | 16 |
| SSLCACertificateFile | `irods_ssl_ca_certificate_file` | `IRODS_SSL_CA_CERTIFICATE_FILE` | |
| SSLCACertificatePath | `irods_ssl_ca_certificate_path` | `IRODS_SSL_CA_CERTIFICATE_PATH` | |
| SSLVerifyServer | `irods_ssl_verify_server` | `IRODS_SSL_VERIFY_SERVER` | hostname |
| SSLCertificateChainFile | `irods_ssl_certificate_chain_file`| `IRODS_SSL_CERTIFICATE_CHAIN_FILE` | |
| SSLCertificateKeyFile | `irods_ssl_certificate_key_file` | `IRODS_SSL_CERTIFICATE_KEY_FILE` | |
| SSLDHParamsFile | `irods_ssl_dh_params_file` | `IRODS_SSL_DH_PARAMS_FILE` | |
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/icommands/configuration.md){target=_blank} (last source update 2025-04-01), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/icommands/data-management/
---
title: "Data management using iCommands"
description: "Navigate collections and list, create, upload, download, move, rename, and remove data in the Data Store with iCommands."
type: Guide
tags:
- iCommands
- Data Management
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/icommands/data_management.md"
title: "CyVerse Learning Materials: docs/ds/icommands/data_management.md"
author: "team:cyverse"
last_modified: "2025-03-25T14:01:09-07:00"
---
# Data management using iCommands
iCommands offers a variety of commands to help you manage your data in the Data Store. In the Data Store, `file` and `directory` are treated as `data objects` and `collections`, respectively. It's perfectly fine to consider these terms interchangeable.
---
## :material-file-arrow-left-right-outline: Display the Current Working Collection
In the Data Store, the **current working collection** is equivalent to the concept of a current working directory in traditional file systems. You can display or change your current working collection using iCommands.
```sh
ipwd
```
By default, after configuring iCommands, your current working collection is set to your **home directory**, which is typically located at:
```sh
//home/
```
> **Note:** Paths in the Data Store always start with the zone name `/iplant`.
---
## :material-file-arrow-left-right-outline: Change the Current Working Collection
1. **Change to a specific collection using an absolute path:**
```sh
icd /iplant/home/myUser/mydata
```
This changes your current working collection to `/iplant/home/myUser/mydata`.
2. **Use a relative path from your current location:**
Assuming your current working collection is `/iplant/home/myUser`:
```sh
icd mydata
```
3. **Return to your home collection:**
```sh
icd "~"
```
> **Note:** The `~` must be quoted to prevent shell expansion by your local shell. Without quotes, it will expand to your local machine's home directory instead of your Data Store home directory.
4. **Move up one level:**
```sh
icd ..
```
---
## :material-file-arrow-left-right-outline: List Data Objects (files) and Collections (directories)
1. **List the content of a collection:**
```sh
ils /iplant/home/myUser/mydata
```
This will display the data objects and collections in the `/iplant/home/myUser/mydata` collection:
```sh
/iplant/home/myUser/mydata:
file1.bin
file2.bin
C- /iplant/home/myUser/mydata/subdir1
```
The `C-` prefix indicates that the item is a collection (directory).
2. **List the content of the current working collection:**
```sh
ils
```
3. **List the contents of a collection in long format with additional details:**
```sh
ils -l /iplant/home/myUser/mydata
```
This command will show the data objects and collections within `/iplant/home/myUser/mydata`, along with their additional details:
```sh
/iplant/home/myUser/mydata:
myUser 0 demoRes1;rs1 436 2024-04-02.13:36 & file1.bin
myUser 1 demoRes2;rs2 436 2024-04-02.13:36 & file1.bin
myUser 0 demoRes1;rs1 700 2024-04-02.17:15 & file2.bin
myUser 1 demoRes2;rs2 700 2024-04-02.17:15 & file2.bin
C- /iplant/home/myUser/mydata/subdir1
```
- Each output line for a data object represents a replica. If iRODS is configured to create multiple replicas, you will see one line for each replica of the data object. For example, if two replicas are created, two lines will be displayed for each file.
- Each line is shown in the `owner replica_id resource_server size creation_time replica_state name` format.
- Possible replica states are:
- `&`: Good
- `X`: Stale
- `?`: Unknown
4. **List the contents of a collection with their access control lists:**
```sh
ils -A /iplant/home/myUser/mydata
```
This command will show the data objects and collections within `/iplant/home/myUser/mydata`, along with their access control lists (ACLs):
```sh
/iplant/home/myUser/mydata:
ACL - g:rodsadmin#iplant:own myUser#iplant:own
Inheritance - Disabled
file1.bin
ACL - g:rodsadmin#iplant:own myUser#iplant:own
file2.bin
ACL - g:rodsadmin#iplant:own myUser#iplant:own
C- /iplant/home/myUser/mydata/subdir1
```
- The `g:` prefix in the ACL username indicates that the user is a group.
- The ACL is displayed in the `username#zone:access_level` format.
- Most common access levels are:
- `read_object`: Allows read access to the data object or collection.
- `modify_object`: Allows modification (write) of the data object or collection.
- `own`: Grants ownership of the data object or collection.
---
## :material-file-arrow-left-right-outline: Make a Collections (directories)
1. **Create a new collection:**
```sh
imkdir /iplant/home/myUser/newCollection
```
2. **Create parent collections if they do not exist:**
```sh
imkdir -p /iplant/home/myUser/parentCollection/newCollection
```
This command creates the `newCollection` along with its parent collection `parentCollection` if it does not already exist.
---
## :material-file-arrow-left-right-outline: Upload Data Objects (files) and Collections (directories) to the Data Store
!!! warning
When uploading your data to the Data Store, avoid using:
- Spaces in names (e.g., `experiment one.fastq`)
- Special characters: \~ \`\` ! @ \# \$ % \^ & \* ( ) + = { } \[ \] \| : ; " ' < \> , ? / and \\
These may cuase issues with Discovery Environment Apps and command-line applications.
**Recommendation:** Use underscores for long names (e.g., `experiment_one.fastq`).
1. **Upload a single file:**
```sh
iput /local/path/file.txt /iplant/home/myUser/
```
This command uploads the file `/local/path/file.txt` to `/iplant/home/myUser/`, creating `/iplant/home/myUser/file.txt` in the Data Store.
2. **Upload a directory and its contents:**
```sh
iput /local/dir /iplant/home/myUser/
```
This command uploads the contents of the directory `/local/dir` to `/iplant/home/myUser/dir` in the Data Store. The uploaded files and subdirectories will be placed within the `/iplant/home/myUser/dir` folder.
3. **Upload with progress output:**
```sh
iput -P /local/path/largefile.dat /iplant/home/myUser/
```
4. **Force upload:**
```sh
iput -f /local/path/largefile.dat /iplant/home/myUser/
```
This command overwrites the existing file in the Data Store without prompting.
5. **Upload and verify checksum:**
```sh
iput -K /local/path/important_data.txt /iplant/home/myUser/
```
This command uploads the file and verifies its integrity by calculating a checksum during transfer.
6. **Upload via resource server:**
```sh
iput -I /local/dir /iplant/home/myUser/
```
This command bypasses the iCAT server for data transfer, directly accessing the specified resource server for optimized performance.
7. **Upload with connection auto-renewal:**
```sh
iput -T /local/dir /iplant/home/myUser/
```
Renews connections every 10 minutes to prevent failures due to connection or firewall issues.
---
## :material-file-arrow-left-right-outline: Download Data Objects (files) and Collections (directories) From the Data Store
1. **Download a data object to a specific local path:**
```sh
iget /iplant/home/myUser/file.txt /local/path/file_new_name.txt
```
This command downloads the data object `/iplant/home/myUser/file.txt` and saves it as `/local/path/file_new_name.txt`.
2. **Download a collection to a specific local path:**
```sh
iget /iplant/home/myUser/dir /local/path/
```
This command downloads the collection `/iplant/home/myUser/dir` and its contents to `/local/path`. A new directory named `dir` will be created under `/local/path`, resulting in `/local/path/dir` containing all the downloaded files and subdirectories.
3. **Download with progress output:**
```sh
iget -P /iplant/home/myUser/largefile.dat /local/dir/
```
4. **Force download:**
```sh
iget -f /iplant/home/myUser/largefile.dat .
```
This command overwrites the local file without prompting if it already exists.
5. **Download and verify checksum:**
```sh
iget -K /iplant/home/myUser/important_data.txt .
```
This command downloads the file and verifies its integrity by calculating the checksum after download and comparing it with the original in the Data Store. This ensures data consistency and detects any corruption during transfer.
6. **Download via resource server:**
```sh
iget -I /iplant/home/myUser/dir /local/dir
```
This command bypasses the iCAT server for data transfer, directly accessing the specified resource server. It optimizes performance for large files by direct connection to the resource server.
7. **Download with connection auto-renewal:**
```sh
iget -T /iplant/home/myUser/dir /local/dir
```
Renews connections every 10 minutes to prevent failures due to connection or firewall issues.
---
## :material-file-arrow-left-right-outline: Remove Data Objects (files) or Collections (directories) From the Data Store
1. **Remove a single data object:**
```sh
irm /iplant/home/myUser/file.txt
```
2. **Remove an empty collection:**
```sh
irmdir /iplant/home/myUser/emptyCollection
```
3. **Remove a collection and its contents recursively:**
```sh
irm -r /iplant/home/myUser/parentCollection
```
4. **Force remove a collection and its contents recursively:**
```sh
irm -rf /iplant/home/myUser/parentCollection
```
---
## :material-file-arrow-left-right-outline: Move/Rename Data Objects (files) or Collections (directories)
1. **Rename a data object:**
```sh
imv /iplant/home/myUser/oldfile.txt /iplant/home/myUser/newfile.txt
```
2. **Move a data object to a different collection:**
```sh
imv /iplant/home/myUser/file.txt /iplant/home/myUser/subcollection/
```
3. **Rename a collection:**
```sh
imv /iplant/home/myUser/oldcollection /iplant/home/myUser/newcollection
```
---
## :material-list-box-outline: Additional Resources
For detailed information on iCommands, refer to the [iRODS iCommands Docs](https://docs.irods.org/4.3.4/icommands/user/){target=_blank}.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/icommands/data_management.md){target=_blank} (last source update 2025-03-25), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/icommands/data-transfer/
---
title: "Data transfer using iCommands"
description: "Download and upload files and folders between a local system and the Data Store with iget and iput, including thread-count tuning."
type: Guide
tags:
- iCommands
- Data Transfer
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/icommands/data_transfer.md"
title: "CyVerse Learning Materials: docs/ds/icommands/data_transfer.md"
author: "team:cyverse"
last_modified: "2025-03-25T14:01:09-07:00"
---
# Data transfer using iCommands
iCommands provides a range of commands designed to efficiently and conveniently transfer large datasets between your local machine and the Data Store.
---
## :material-package-variant-closed: Download Data from the Data Store
1. **Download a collection to a specific local path:**
```sh
iget -P -f -K -T /iplant/home/myUser/mydir /local/dir/
```
This command downloads the collection `/iplant/home/myUser/mydir` and its contents to `/local/dir/`. A new directory named `mydir` will be created under `/local/dir/`, resulting in `/local/dir/mydir` containing all the downloaded files and subdirectories.
- The `-P` flag outputs download progress.
- The `-f` flag forces overwriting of existing local files without confirmation.
- The `-K` flag ensures file integrity by verifying checksums after the download.
- The `-T` flag renews connections every 10 minutes to prevent failures due to connection or firewall issues.
2. **Sync a collection in the Data Store with a local directory:**
```sh
irsync -r -K i:/iplant/home/myUser/mydir /local/dir/
```
This command downloads the collection `/iplant/home/myUser/mydir` and its contents to `/local/dir/`, syncing any new updates made to the original data.
3. **Download a collection using multiple parallel transfers:**
```sh
iget -P -f -K -T -N 10 /iplant/home/myUser/mydir /local/dir/
```
This command downloads the collection `/iplant/home/myUser/mydir` and its contents to `/local/dir/` using 10 parallel transfer threads. Utilizing more transfer threads can maximize I/O and network bandwidth, often speeding up the transfer process significantly.
- The `-N` flag sets the maximum number of threads to use for the transfer.
!!! warning "Thread Count Considerations"
1. Higher thread counts require more CPU and memory. Excessive threads may overload your system, causing performance issues. For example, Discovery Environment (DE) apps limit transfer threads to 4 due to RAM constraints.
2. The Data Store limits concurrent connections, potentially restricting high thread counts.
---
## :material-package-variant-closed: Upload Data to the Data Store
1. **Upload a local directory to a specific path in the Data Store:**
```sh
iput -P -f -K -T /local/dir/ /iplant/home/myUser/mydir/
```
This command uploads the local directory `/local/dir/` and its contents to the Data Store at `/iplant/home/myUser/mydir/`. A new directory named `dir` will be created under `/iplant/home/myUser/mydir/`, resulting in `/iplant/home/myUser/mydir/dir` containing all the uploaded files and subdirectories.
- The `-P` flag outputs upload progress.
- The `-f` flag forces overwriting of existing files in the Data Store without confirmation.
- The `-K` flag ensures file integrity by verifying checksums after upload.
- The `-T` flag renews connections every 10 minutes to prevent failures due to connection or firewall issues.
2. **Sync a local directory with a collection in the Data Store:**
```sh
irsync -r -K /local/dir/ i:/iplant/home/myUser/mydir/
```
This command uploads the local directory `/local/dir/` to `/iplant/home/myUser/mydir/` in the Data Store, syncing any new or updated files.
3. **Upload a local directory using multiple parallel transfers:**
```sh
iput -P -f -K -T -N 10 /local/dir/ /iplant/home/myUser/mydir/
```
This command uploads the local directory `/local/dir/` and its contents to `/iplant/home/myUser/mydir/` in the Data Store using 10 parallel transfer threads. This can significantly speed up the upload process by maximizing I/O and network bandwidth.
- The `-N` flag sets the maximum number of threads to use for the transfer.
!!! warning "Thread Count Considerations"
1. Higher thread counts require more CPU and memory. Excessive threads may overload your system, causing performance issues. For example, Discovery Environment (DE) apps limit transfer threads to 4 due to RAM constraints.
2. The Data Store limits concurrent connections, potentially restricting high thread counts.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/icommands/data_transfer.md){target=_blank} (last source update 2025-03-25), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/icommands/access-management/
---
title: "Access management using iCommands"
description: "List and change user and group access permissions on Data Store data objects and collections with iCommands."
type: Guide
tags:
- iCommands
- Permissions
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/icommands/access_management.md"
title: "CyVerse Learning Materials: docs/ds/icommands/access_management.md"
author: "team:cyverse"
last_modified: "2025-03-25T14:01:09-07:00"
---
# Access management using iCommands
iCommands provides features to manage access of users and groups to data in the Data Store. The `ichmod` command allows users to manage access permissions for data objects (files) and collections (directories). The `ils -A` command displays the current access levels assigned to users for a given file or directory.
---
## :material-security: List Access Permissions of Users for a Data Object or Collection
```sh
ils -A
```
The `-A` flag in `ils` command displays access permissions in the result.
This command will show the data objects and collections along with their access control lists (ACLs). For example:
```sh
/iplant/home/myUser/mydata:
ACL - g:rodsadmin#iplant:own myUser#iplant:own
Inheritance - Disabled
file1.bin
ACL - g:rodsadmin#iplant:own myUser#iplant:own
file2.bin
ACL - g:rodsadmin#iplant:own myUser#iplant:own
C- /iplant/home/myUser/mydata/subdir1
```
- The `g:` prefix in the ACL username indicates that the user is a group.
- The ACL is displayed in the `username#zone:access_level` format.
- Most common access levels are:
- `read_object`: Allows read access to the data object or collection.
- `modify_object`: Allows modification (write) of the data object or collection.
- `own`: Grants ownership of the data object or collection.
### Example Usage
1. **List current access levels for a data object or collection:**
```sh
ils -A /iplant/home/myUser/mydata
```
## :material-security: Change a User's or Group's Access Permission for a Data Object or Collection
```sh
ichmod
```
### Access Levels
| Access Level | Description |
|-------------|-------------|
| `null` | Removes all permissions |
| `read` | Allows reading the object or collection |
| `write` | Allows reading and modifying the object or collection |
| `own` | Grants full control, including the ability to change permissions |
### Example Usage
1. **Grant a user read permission to a data object:**
```sh
ichmod read anotherUser /iplant/home/myUser/file.txt
```
2. **Grant a user from a different zone read permission to a data object:**
```sh
ichmod read anotherUser#anotherZone /iplant/home/myUser/file.txt
```
3. **Grant a user read permission to a collection and its contents:**
```sh
ichmod -r read anotherUser /iplant/home/myUser/dir
```
4. **Grant a user write permission to a collection and its contents:**
```sh
ichmod -r write anotherUser /iplant/home/myUser/dir
```
5. **Grant a user owner permission to a collection and its contents:**
```sh
ichmod -r owner anotherUser /iplant/home/myUser/dir
```
6. **Remove access permission from a user to a collection and its contents:**
```sh
ichmod -r none anotherUser /iplant/home/myUser/dir
```
---
## :material-security: Enable or Disable Access Permission Inheritance for a Collection
When inheritance is enabled for a collection, any new data objects or subcollections created within it will automatically inherit the same access permissions as the parent collection.
```sh
ichmod
```
### Inheritance Options
| Flag | Description |
|------|-------------|
| `inherit` | Enable access inheritance. Data objects and sub-collections inherit permissions from the parent collection |
| `noinherit` | Disable access inheritance. Data objects and sub-collections do not inherit permissions from the parent collection |
### Example Usage
1. **Enable inheritance for a collection:**
```sh
ichmod inherit /iplant/home/myUser/dir
```
2. **Disable inheritance for a collection:**
```sh
ichmod noinherit /iplant/home/myUser/dir
```
3. **Enable inheritance recursively for a collection and its subcollections:**
```sh
ichmod -r inherit /iplant/home/myUser/dir
```
4. **Disable inheritance recursively for a collection and its subcollections:**
```sh
ichmod -r noinherit /iplant/home/myUser/dir
```
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/icommands/access_management.md){target=_blank} (last source update 2025-03-25), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/icommands/metadata-management/
---
title: "Metadata management using iCommands"
description: "List, add, and remove metadata on Data Store data objects, collections, resources, and users with iCommands."
type: Guide
tags:
- iCommands
- Metadata
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/icommands/metadata_management.md"
title: "CyVerse Learning Materials: docs/ds/icommands/metadata_management.md"
author: "team:cyverse"
last_modified: "2025-03-25T14:01:09-07:00"
---
# Metadata management using iCommands
GoCommands provides features to manage metadata for data objects, collections, resources, and users in the Data Store using the `imeta` command.
Metadata consists of three components:
- **Name (Attribute):** The name of information.
- **Value:** The actual data or information.
- **Unit (Optional):** Specifies the unit of measurement, if applicable.
---
## :material-tag-edit-outline: List Metadata of Data Objects, Collections, Resources, or Users
```sh
imeta ls [flags] ...
```
### iRODS Objects
| iROD Object | Flag | Description |
|-------------|-------------|--------|
| `data object` | `-d` | List metadata of data objects |
| `collection` | `-C` | List metadata of collections |
| `resource` | `-R` | List metadata of resources |
| `user` | `-u` | List metadata of users |
### Example Usage
1. **List metadata of a data object:**
```sh
imeta ls -d /iplant/home/myUser/file.txt
```
2. **List metadata of a collection:**
```sh
imeta ls -C /iplant/home/myUser/dir
```
3. **List metadata of a resource:**
```sh
imeta ls -R myResc
```
4. **List metadata of a user:**
```sh
imeta ls -u myUser
```
## :material-tag-edit-outline: Add Metadata to Data Objects, Collections, Resources, or Users
```sh
imeta add [flags] [metadata-unit]
```
### iRODS Objects
| iROD Object | Flag | Description |
|-------------|-------------|--------|
| `data object` | `-d` | List metadata of data objects |
| `collection` | `-C` | List metadata of collections |
| `resource` | `-R` | List metadata of resources |
| `user` | `-u` | List metadata of users |
### Example Usage
1. **Add metadata to a data object:**
```sh
imeta add -d /iplant/home/myUser/file.txt meta_name meta_value
```
1. **Add metadata to a data object with metadata-unit:**
```sh
imeta add -d /iplant/home/myUser/file.txt meta_name meta_value meta_unit
```
3. **Add metadata to a collection:**
```sh
imeta add -C /iplant/home/myUser/dir meta_name meta_value
```
4. **Add metadata to a resource:**
```sh
imeta add -R myResc meta_name meta_value
```
5. **Add metadata to a user:**
```sh
imeta add -u myUser meta_name meta_value
```
---
## :material-tag-edit-outline: Remove Metadata from Data Objects, Collections, Resources, or Users
Remove metadata by name (attribute) and value:
```sh
imeta rm [flags]
```
Remove metadata by ID:
```sh
imeta rmi [flags]
```
**Note:** The `metadata-ID` is a numeric identifier for the metadata. It can be obtained from the output of the `imeta ls` command.
### iRODS Objects
| iROD Object | Flag | Description |
|-------------|-------------|--------|
| `data object` | `-d` | List metadata of data objects |
| `collection` | `-C` | List metadata of collections |
| `resource` | `-R` | List metadata of resources |
| `user` | `-u` | List metadata of users |
### Example Usage
1. **Remove metadata from a data object by name:**
```sh
imeta rm -d /iplant/home/myUser/file.txt meta_name meta_value
```
2. **Remove metadata from a data object by ID:**
```sh
imeta rmi -d /iplant/home/myUser/file.txt 979206950
```
3. **Remove metadata from a collection:**
```sh
imeta rm -C /iplant/home/myUser/dir meta_name meta_value
```
4. **Remove metadata from a resource:**
```sh
imeta rm -R myResc meta_name meta_value
```
5. **Remove metadata from a user:**
```sh
imeta rm -u myUser meta_name meta_value
```
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/icommands/metadata_management.md){target=_blank} (last source update 2025-03-25), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/icommands/troubleshooting/
---
title: "iCommands troubleshooting and issue report"
description: "Diagnose common iCommands problems and report bugs to the iRODS developers."
type: Guide
tags:
- iCommands
- Troubleshooting
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/icommands/issue_report.md"
title: "CyVerse Learning Materials: docs/ds/icommands/issue_report.md"
author: "team:cyverse"
last_modified: "2025-03-25T14:01:09-07:00"
---
# iCommands troubleshooting and issue report
## :material-bullseye-arrow: Troubleshooting
### Installing iCommands on Windows and macOS
iCommands has limited support for Windows and macOS. While there are options available, such as using Windows Subsystem for Linux (WSL) or custom iCommands packages for macOS, we highly recommend using GoCommands as a more versatile alternative.
GoCommands is a cross-platform tool that offers similar functionality to iCommands with several advantages:
- **Cross-platform compatibility**: Works seamlessly on Windows, macOS, and Linux
- **No installation required**: Simply download and run the executable
- **User-friendly commands**: Provides essential data access functions
For detailed instructions on downloading, setting up, and using GoCommands, please refer to the [GoCommands documentation](https://unm-carc.github.io/cyverse/data-store/gocommands/overview/).
### Running iCommands in an HPC Environment
To use pre-installed iCommands in an HPC environment:
```sh
module load irods
```
This command provides access to iRODS iCommands.
### Request Support
If you encounter an issue that you cannot resolve, please contact [support@cyverse.org](mailto:support@cyverse.org) for assistance. Your Data Store access via iCommands may be limited or fail due to various factors, including configuration issues, network problems, authentication errors, or data policies. The support team is available to help you identify and resolve these issues.
---
## :material-bug-check-outline: Report Bugs
Encountered a bug in iCommands? We encourage you to report it on the iCommands [GitHub issues page](https://github.com/irods/irods_client_icommands/issues){target=_blank}. Your detailed bug reports are invaluable for improving the stability and usability of iCommands for the entire iRODS community. When submitting, please provide as much information as possible to help with diagnosis and resolution.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/icommands/issue_report.md){target=_blank} (last source update 2025-03-25), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/sftp/overview/
---
title: "Manage your data using SFTP"
description: "SFTP access to the Data Store: when to use it, its limits for large files, and the SFTPGo server behind it."
type: Guide
tags:
- SFTP
- Data Store
- Data Transfer
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/sftp/index.md"
title: "CyVerse Learning Materials: docs/ds/sftp/index.md"
author: "team:cyverse"
last_modified: "2025-09-19T11:30:27-07:00"
---
# Manage your data using SFTP
SFTP (Secure File Transfer Protocol) is a widely adopted network protocol for secure file transfer and management. It operates over an encrypted communication channel, ensuring safe data exchange between the client and server. With broad compatibility across various environments, SFTP offers a flexible and reliable solution for managing data.
This guide covers how to configure SFTP clients to efficiently manage your data in the Data Store.
---
## Limitations
**Using SFTP for File Transfers**
SFTP is ideal for transferring small files or small collections of files. While there is no strict size limit, it is not recommended for files larger than 10 GiB due to performance issues.
**Alternatives for Large Files**
For large files or extensive datasets, consider using **GoCommands** or **iCommands** instead. These tools offer better performance and efficiency for handling large data transfers.
For more details on GoCommands and iCommands, visit their respective documentation pages:
- [GoCommands](https://unm-carc.github.io/cyverse/data-store/gocommands/overview/)
- [iCommands](https://unm-carc.github.io/cyverse/data-store/icommands/overview/)
The [SFTP contents](https://unm-carc.github.io/cyverse/data-store/sftp/) page lists the client guides and configuration pages in this section.
---
## Acknowledgments
The SFTP functionality for the Data Store is powered by [SFTPGo](https://github.com/drakkan/sftpgo){target=_blank}, an open-source, fully featured, and highly configurable SFTP server created by Nicola Murino. SFTPGo supports various storage backends, including local filesystems, S3 Object Storage, Google Cloud Storage, and Azure Blob Storage. CyVerse extended SFTPGo's capabilities by implementing a new backend module specifically for iRODS, enabling SFTP access to the Data Store. We extend our gratitude to Nicola Murino and the SFTPGo project for making this integration possible.
!!! info "Under the hood"
SFTP access is served by SFTPGo with a custom iRODS storage backend, behind the same `data.cyverse.org` entry point as the other Data Store services; see [Data Store](https://docs.cyverse.org/platform/data-store/){target=_blank} and [Data Store administration](https://docs.cyverse.org/operations/data-store/){target=_blank} in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/sftp/index.md){target=_blank} (last source update 2025-09-19), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/sftp/cli/
---
title: "SFTP access using command-line tools"
description: "Connect to the Data Store with your operating system's built-in sftp client and transfer files from the terminal."
type: Guide
tags:
- SFTP
- Command Line
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/sftp/cli.md"
title: "CyVerse Learning Materials: docs/ds/sftp/cli.md"
author: "team:cyverse"
last_modified: "2025-04-01T14:25:43-07:00"
---
# SFTP access using command-line tools
Most operating systems include built-in SFTP clients, enabling command-line access to the Data Store. This guide outlines the basics of using SFTP tools on Linux, macOS, and Windows.
!!! tip "Windows Users"
Windows 10 and later versions include an SFTP client. For earlier versions, you may need to use a third-party command-line tool like `WinSCP` or a GUI tool like `FileZilla`.
## :material-account-circle-outline: SFTP Access Information
Use the following credentials to connect to the Data Store:
| Key | Value |
|-----------------|----------------------|
| `hostname` | `data.cyverse.org` |
| `port` | `22` |
| `username` | `` |
| `password` | `` |
Use these credentials for anonymous access to the Data Store:
| Key | Value |
|-------------------|-------|
| `username` | `anonymous` |
| `password` | (leave empty) |
---
## :material-console: Connect to the Data Store
To connect using SFTP, open a terminal and run:
```sh
sftp @data.cyverse.org
```
Upon successful connection, you'll see a prompt like this:
```sh
$ sftp @data.cyverse.org
Connected to data.cyverse.org.
sftp>
```
> **Note:** Output may vary depending on the operating system. The example above is from Linux.
---
## :material-console: Basic SFTP Commands
Once connected, you can use these common SFTP commands:
- `ls`: List files and directories
- `cd`: Change directory
- `pwd`: Display current directory
- `get`: Download a file from the Data Store
- `put`: Upload a file to the Data Store
- `mkdir`: Create a directory
- `rmdir`: Remove an empty directory
- `rm`: Delete a file
To close the SFTP connection, use the `exit` or `bye` command.
Use the `help` or `?` command to see a list of available SFTP commands.
---
## :material-folder-multiple-outline: Top-level Directories
Once connected, you will see two directories in the root:
- ``: Your home directory (`/iplant/home/` in the Data Store). You have read and write permissions. Note that anonymous users do not have a home directory.
- `shared`: Community-shared data directory (`/iplant/home/shared` in the Data Store). You have only read permission.
> **Note:** A `.ssh` directory may appear in the root, but it is not writable. This directory is distinct from the `//.ssh` directory and should be ignored.
---
## :material-console: Examples
1. **List files in your home directory:**
```sh
ls /myUser
```
2. **Download a file:**
```sh
get /myUser/myfile.txt
```
3. **Upload a file:**
```sh
put localfile.txt /myUser/
```
4. **Create a new directory:**
```sh
mkdir /myUser/newdir
```
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/sftp/cli.md){target=_blank} (last source update 2025-04-01), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/sftp/filezilla/
---
title: "SFTP access using FileZilla"
description: "Install FileZilla and connect to the Data Store over SFTP to transfer files with a graphical client."
type: Guide
tags:
- SFTP
- FileZilla
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/sftp/filezilla.md"
title: "CyVerse Learning Materials: docs/ds/sftp/filezilla.md"
author: "team:cyverse"
last_modified: "2025-03-27T13:28:03-07:00"
---
# SFTP access using FileZilla
FileZilla is a free and open-source, cross-platform GUI FTP software, consisting of FileZilla Client and FileZilla Server. The FileZilla Client is available for Windows, Linux, and macOS, allowing you to access the Data Store.
## :material-cog-outline: Installation
To install the FileZilla Client, follow these steps:
1. **Download FileZilla Client:**
Visit the [FileZilla website](https://filezilla-project.org/download.php?type=client){target=_blank} and download the client version suitable for your operating system (Windows, Linux, or macOS).
2. **Install FileZilla:**
- **Windows:** Double-click the downloaded `.exe` file and follow the installation wizard.
- **macOS:** Open the downloaded package and drag the FileZilla application to your Applications folder.
- **Linux:** Use your distribution's package manager to install FileZilla. Alternatively, you can compile it from source if necessary.
3. **Launch FileZilla:**
After installation, launch the FileZilla Client to start using it.
---
## :simple-filezilla: Connect to the Data Store
In the FileZilla window, fill in the following fields:
- **Host:** `data.cyverse.org`
- **Username:** ``
- **Password:** ``
- **Port:** `22`
Use these credentials for anonymous access to the Data Store:
- **Username:** `anonymous`
- **Password:** (leave empty)
{ width="600" }
Click the **Quickconnect** button to establish the connection.
---
## :simple-filezilla: Basic Usage
{ width="600" }
The FileZilla interface is divided into two main sections:
- **Left pannel:** Show data on your local machine
- **Right pannel:** Display data in the Data Store
**To navigate:**
- Click on directory names to move in and out of folders
**To transfer files:**
1. Select the desired files or directories
2. Drag them to the target directory in the opposite panel
3. Drop to initiate the transfer
This drag-and-drop functionality allows for easy file movement between your local system and the Data Store.
---
## :material-folder-multiple-outline: Top-level Directories
Once connected, you will see two directories in the root:
- ``: Your home directory (`/iplant/home/` in the Data Store). You have read and write permissions. Note that anonymous users do not have a home directory.
- `shared`: Community-shared data directory (`/iplant/home/shared` in the Data Store). You have only read permission.
> **Note:** A `.ssh` directory may appear in the root, but it is not writable. This directory is distinct from the `//.ssh` directory and should be ignored.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/sftp/filezilla.md){target=_blank} (last source update 2025-03-27), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/sftp/cyberduck/
---
title: "SFTP access using Cyberduck"
description: "Install Cyberduck and connect to the Data Store over SFTP to transfer files with a graphical client."
type: Guide
tags:
- SFTP
- Cyberduck
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/sftp/cyberduck.md"
title: "CyVerse Learning Materials: docs/ds/sftp/cyberduck.md"
author: "team:cyverse"
last_modified: "2025-03-27T13:28:03-07:00"
---
# SFTP access using Cyberduck
Cyberduck is a free, open-source GUI FTP client available for both Windows and macOS. It supports various protocols such as FTP, SFTP, and WebDAV, allowing you to access the Data Store and other cloud storage services like Amazon S3, Google Drive, and Dropbox.
This guide demonstrates how to use Cyberduck for SFTP access to the Data Store.
## :material-cog-outline: Installation
To install Cyberduck, follow these steps:
1. **Download Cyberduck:**
Visit the [Cyberduck website](https://cyberduck.io/download/){target=_blank} and download the appropriate version for your operating system (Windows or macOS).
2. **Install Cyberduck:**
- **Windows:** Double-click the downloaded `.exe` file and follow the installation wizard.
- **macOS:** Double-click the downloaded `.zip` file to extract the contents. Open the extracted folder and find the `Cyberduck.app` icon. Drag and drop this icon into your Applications folder.
3. **Launch Cyberduck:**
After installation, launch the Cyberduck to start using it.
---
## :material-duck: Connect to the Data Store
{ width="600" }
In the Cyberduck window, click **Open Connection** button to create a new profile.
{ width="600" }
In the popup window, fill in the following fields:
- **Protocol:** `SFTP (SSH File Transfer Protocol)`
- **Server:** `data.cyverse.org`
- **Port:** `22`
- **Username:** ``
- **Password:** ``
Use these credentials for anonymous access to the Data Store:
- **Username:** `anonymous`
- **Password:** (leave empty)
Click the **Connect** button to establish the connection.
---
## :material-duck: Basic Usage
**To navigate:**
{ width="600" }
- Click on directory names to move into folders
- Click the top drop-down box and click directory name to move out of folders
**To transfer files:**
{ width="600" }
1. Select the desired files or directories
2. Right-click on the selected items and choose `Download To...` from the context menu
> **Note:** Additionally, you can utilize the drag-and-drop feature to easily upload files and directories from your local system to the Data Store, and vice versa.
---
## :material-folder-multiple-outline: Top-level Directories
Once connected, you will see two directories in the root:
- ``: Your home directory (`/iplant/home/` in the Data Store). You have read and write permissions. Note that anonymous users do not have a home directory.
- `shared`: Community-shared data directory (`/iplant/home/shared` in the Data Store). You have only read permission.
> **Note:** A `.ssh` directory may appear in the root, but it is not writable. This directory is distinct from the `//.ssh` directory and should be ignored.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/sftp/cyberduck.md){target=_blank} (last source update 2025-03-27), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/sftp/directory-structure/
---
title: "SFTP directory structure"
description: "How your home directory and community-shared Data Store directories appear when you connect over SFTP."
type: Reference
tags:
- SFTP
- Data Store
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/sftp/structure.md"
title: "CyVerse Learning Materials: docs/ds/sftp/structure.md"
author: "team:cyverse"
last_modified: "2025-03-27T12:34:29-07:00"
---
# SFTP directory structure
When accessing the Data Store via SFTP, you'll encounter a different directory structure compared to other tools like GoCommands and iCommands. SFTP users have access to following directories:
## :material-folder-multiple-outline: Home Directory
`/`
- Maps to your iRODS home directory `/iplant/home/`
- Read and write access for the owner
- Not provided to anonymous users
## :material-folder-multiple-outline: Community-shared Data Directory
`/shared`
- Maps to the iRODS community-shared data directory: `/iplant/home/shared`
- Read-only access
## :material-folder-multiple-outline: `.ssh` Directory
`/.ssh`
- Automatically generated by the SFTP service
- Not used and should be ignored
- Distinct from `//.ssh`
For password-less public-key authentication, use `//.ssh` instead.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/sftp/structure.md){target=_blank} (last source update 2025-03-27), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/sftp/public-key-auth/
---
title: "SFTP public-key authentication configuration"
description: "Set up SSH public-key authentication for password-less SFTP access to the Data Store."
type: Guide
tags:
- SFTP
- SSH Keys
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/sftp/public_key_auth.md"
title: "CyVerse Learning Materials: docs/ds/sftp/public_key_auth.md"
author: "team:cyverse"
last_modified: "2025-03-26T12:33:22-07:00"
---
# SFTP public-key authentication configuration
Set up password-less authentication for SFTP access by uploading your SSH public key to the Data Store.
## :material-account-key-outline: Setting Up Public-key Authentication
1. **Generate an SSH key (if needed):**
```sh
ssh-keygen -t rsa -b 4096 -C "your_email@example.com"
```
This creates a private key `id_rsa` and a public key `id_rsa.pub` in the `~/.ssh` directory.
2. **Connect to the Data Store via SFTP:**
```sh
sftp myUser@data.cyverse.org
```
3. **Create the `.ssh` directory in your Data Store home:**
In the SFTP prompt, run:
```sh
mkdir /myUser/.ssh
```
This creates the `.ssh` directory at `/iplant/home//.ssh` in the Data Store.
> **Note:** A `.ssh` directory may appear in the root (`/.ssh`), but it is not writable. This directory is distinct from the `//.ssh` directory and should be ignored.
4. **Copy your SSH public key to the Data Store:**
Still in the SFTP prompt, copy your local `~/.ssh/id_rsa.pub` file to the Data Store:
```sh
put ~/.ssh/id_rsa.pub /myUser/.ssh/authorized_keys
```
This registers your public key for password-less authentication.
4. **Exit and reconnect to verify:**
```sh
quit
sftp myUser@data.cyverse.org
```
It should not ask for a password this time.
---
## :material-account-key-outline: Advanced Configuration
For advanced usage, you can control public-key access by manually editing the `/iplant/home//.ssh/authorized_keys` file in the Data Store. This process involves downloading the file, making changes, and then uploading it back. Here's how to do it:
1. **Download the `authorized_keys` File:**
Connect to the Data Store via SFTP and download the file:
```sh
sftp myUser@data.cyverse.org
get /myUser/.ssh/authorized_keys
quit
```
2. **Edit the file locally with a editor:**
Open it with a text editor (e.g., `vi`, `nano`):
```sh
vi authorized_keys
```
Add parameters in `key=value` format before each SSH key. Example:
```sh
expiry-time="20250320" from="10.11.12.13" ssh-rsa AAAAB3Nza... myUser
```
3. **Upload the modified file back to the Data Store:**
Reconnect via SFTP and upload:
```sh
sftp myUser@data.cyverse.org
put authorized_keys /myUser/.ssh/
quit
```
> **Note:** Configuration changes are only applied during user authentication. Therefore, modifications do not affect users or clients that are already logged in.
### Available Parameters
| Parameter | Description | Example |
|-------------|-------------|---------|
| `expiry-time` | Sets expiration date-time in `YYYYMMDD`, `YYYYMMDDhhmm`, or `YYYYMMDDhhmmss` format | `expiry-time="20250320"` |
| `from` | Allows access from specific IP addresses. Use IP address, CIDR, or `!` prefix to negate. Separate multiple entries with commas | `from="10.11.12.13,!10.11.12.14"` |
| `home` | Sets a specific home collection path for SFTP access within the Data Store. Use absolute path of the collection in the Data Store | `home=/iplant/home/myUser/sftp_home` |
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/sftp/public_key_auth.md){target=_blank} (last source update 2025-03-26), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/webdav/overview/
---
title: "Manage your data using WebDAV"
description: "WebDAV (HTTPS) access to the Data Store at data.cyverse.org: when to use it and its limits for large files."
type: Guide
tags:
- WebDAV
- Data Store
- Data Transfer
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/webdav/index.md"
title: "CyVerse Learning Materials: docs/ds/webdav/index.md"
author: "team:cyverse"
last_modified: "2025-03-27T13:53:00-07:00"
---
# Manage your data using WebDAV
WebDAV (Web Distributed Authoring and Versioning) is a network protocol built on top of HTTP/HTTPS, enabling users to manage data on web servers. It operates over HTTP/HTTPS, ensuring secure data exchange between the client and server, with broad compatibility across various environments. WebDAV provides a flexible and reliable solution for managing data.
This guide covers how to use WebDAV to efficiently manage your data in the Data Store.
---
## Limitations
**Using WebDAV for File Transfers**
WebDAV is ideal for transferring small files or small collections of files. While there is no strict size limit, it is not recommended for files larger than 10 GiB due to performance issues.
**Alternatives for Large Files**
For large files or extensive datasets, consider using **GoCommands** or **iCommands** instead. These tools offer better performance and efficiency for handling large data transfers.
For more details on GoCommands and iCommands, visit their respective documentation pages:
- [GoCommands](https://unm-carc.github.io/cyverse/data-store/gocommands/overview/)
- [iCommands](https://unm-carc.github.io/cyverse/data-store/icommands/overview/)
The [WebDAV contents](https://unm-carc.github.io/cyverse/data-store/webdav/) page lists the client guides and data locations in this section.
!!! info "Under the hood"
WebDAV access is served by Apache with the davrods module (with a Varnish cache) at `data.cyverse.org`; see [Data Store](https://docs.cyverse.org/platform/data-store/){target=_blank} and [Data Store administration](https://docs.cyverse.org/operations/data-store/){target=_blank} in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/webdav/index.md){target=_blank} (last source update 2025-03-27), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/webdav/cli/
---
title: "WebDAV access using command-line tools"
description: "Connect to and transfer Data Store files over WebDAV from the terminal, with worked command-line examples."
type: Guide
tags:
- WebDAV
- Command Line
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/webdav/cli.md"
title: "CyVerse Learning Materials: docs/ds/webdav/cli.md"
author: "team:cyverse"
last_modified: "2025-03-27T13:53:00-07:00"
---
# WebDAV access using command-line tools
You can access the Data Store via command-line tools using WebDAV, which is especially useful for integrating data into analysis pipelines or scripts. This guide covers using `curl`, a widely used command-line tool, for efficient WebDAV access and data management.
## :material-account-circle-outline: WebDAV Access Information
Use the following credentials to connect to the Data Store:
| Key | Value |
|-----------------|----------------------|
| `URL` | `https://data.cyverse.org/dav` |
| `username` | `` |
| `password` | `` |
Use these credentials for anonymous access to the Data Store:
| Key | Value |
|-------------------|-------|
| `username` | `anonymous` |
| `password` | (leave empty) |
---
## :material-console: Connect to the Data Store
To list the contents of your home directory in the Data Store using WebDAV, open a terminal and run the following command:
```sh
curl --user ':' https://data.cyverse.org/dav/iplant/home/
```
- Use your CyVerse username and password with `--user` flag to login
- `https://data.cyverse.org/dav` is the URL of the WebDAV root. Add Data Store path after for URL
- `--user ':'`: This option provides your CyVerse username and password for authentication. Replace and with your actual credentials.
- `https://data.cyverse.org/dav/iplant/home/`: This is the URL of your home directory within the Data Store's WebDAV server. `https://data.cyverse.org/dav` is the root WebDAV URL, and `/iplant/home/` specifies the path to your home directory in the Data Store.
The output will be an HTML-formatted text displaying files and directories in a table layout.
---
## :material-web: Data Locations
1. **User Data**
Users can access their data at:
```sh
https://data.cyverse.org/dav/iplant/home//
```
2. **Public Data (Read-Only Access)**
Anonymous users can access public data at:
```sh
https://data.cyverse.org/dav-anon/
```
3. **Community/Project Data**
To access project-specific data stored in iRODS at `/iplant/home/shared//`, use:
```sh
https://data.cyverse.org/dav/iplant/projects//
```
4. **CyVerse Curated Data (DOI-Backed Datasets)**
Access curated datasets with DOIs in the Data Commons at:
```sh
https://data.cyverse.org/dav-anon/iplant/commons/cyverse_curated/
```
---
## :material-console: Examples
1. **List files in your home directory:**
```sh
curl https://data.cyverse.org/dav/iplant/home/myUser
```
This prints files and directories in HTML format.
2. **Read a file:**
```sh
curl https://data.cyverse.org/dav/iplant/home/myUser/myfile.txt
```
This displays the content of `myfile.txt` in the terminal.
3. **Download a file:**
```sh
curl -O https://data.cyverse.org/dav/iplant/home/myUser/myfile.txt
```
This saves `myfile.txt` to the current directory.
```sh
curl -o newfile.txt https://data.cyverse.org/dav/iplant/home/myUser/myfile.txt
```
This saves `myfile.txt` as `newfile.txt` in the current directory.
4. **Upload a file:**
```sh
curl -T localfile.txt https://data.cyverse.org/dav/iplant/home/myUser
```
2. **Creat a new directory:**
```sh
curl -X MKCOL https://data.cyverse.org/dav/iplant/home/myUser/newdir
```
4. **Renaming a File**:
```
curl -X MOVE --header "Destination: https://data.cyverse.org/dav/iplant/home/myUser/new_name.txt" https://data.cyverse.org/dav/iplant/home/myUser/old_name.txt
```
This renames `old_name.txt` to `new_name.txt`.
5. **Deleting Files/Folders**:
```
curl -X DELETE https://data.cyverse.org/dav/iplant/home/myUser/file_to_delete.txt
```
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/webdav/cli.md){target=_blank} (last source update 2025-03-27), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/webdav/web-browsers/
---
title: "WebDAV access using web browsers"
description: "Browse and download Data Store files over WebDAV in any web browser."
type: Guide
tags:
- WebDAV
- Web Browser
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/webdav/browser.md"
title: "CyVerse Learning Materials: docs/ds/webdav/browser.md"
author: "team:cyverse"
last_modified: "2025-03-27T13:53:00-07:00"
---
# WebDAV access using web browsers
WebDAV, an extension of HTTP, allows browsing and downloading data directly through any web browser.
## :material-google-chrome: Connect to the Data Store
In your web browser, enter the URL of the data location.
{ width="600" }
If the data requires authentication, you will be prompted to log in:
- **Username:** ``
- **Password:** ``
For anonymous access:
- **Username:** `anonymous`
- **Password:** (leave empty)
{ width="300" }
Web browsers allow you to list directories, view text file contents, and download files. To manage data fully (e.g., upload, move, or delete files), use **GoCommands**, **iCommands**, **SFTP**, or **WebDAV Command-line Tools** instead.
---
## :material-web: Data Locations
1. **User Data**
Users can access their data at:
```sh
https://data.cyverse.org/dav/iplant/home//
```
2. **Public Data (Read-Only Access)**
Anonymous users can access public data at:
```sh
https://data.cyverse.org/dav-anon/
```
3. **Community/Project Data**
To access project-specific data stored in iRODS at `/iplant/home/shared//`, use:
```sh
https://data.cyverse.org/dav/iplant/projects//
```
4. **CyVerse Curated Data (DOI-Backed Datasets)**
Access curated datasets with DOIs in the Data Commons at:
```sh
https://data.cyverse.org/dav-anon/iplant/commons/cyverse_curated/
```
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/webdav/browser.md){target=_blank} (last source update 2025-03-27), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/webdav/file-managers/
---
title: "WebDAV access using file managers"
description: "Mount the Data Store in macOS Finder, Windows File Explorer, GNOME Files, or a Linux terminal over WebDAV."
type: Guide
tags:
- WebDAV
- File Manager
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/webdav/file_manager.md"
title: "CyVerse Learning Materials: docs/ds/webdav/file_manager.md"
author: "team:cyverse"
last_modified: "2025-03-27T13:53:00-07:00"
---
# WebDAV access using file managers
Most operating systems include a file manager that supports WebDAV, allowing you to mount a WebDAV directory as part of the filesystem. This enables other applications on the same computer to access data in the Data Store as if it were local.
### :simple-apple: macOS Finder
Follow these steps to connect using macOS Finder:
1. Open Finder.
2. In the menu bar, click **Go** → **Connect to Server** (*Command + K*).
3. Enter the WebDAV URL for the folder you want to access.
4. If prompted, enter your **CyVerse username** and **password**.
---
### :material-microsoft-windows-classic: Windows File Exploer
To connect using Windows File Explorer:
1. Open File Explorer.
2. Right-click **This PC** and select **Map Network Drive**.
3. Click **Choose a custom network location**, then **Next**.
4. Enter the WebDAV URL for the folder you want to access.
5. If prompted, enter your **CyVerse username** and **password**.
---
### :material-gnome: Gnome Files
To connect using GNOME Files on Linux:
1. Open **Files**.
2. In the sidebar, click **Other Locations**.
3. Under **Connect to Server**, enter the WebDAV URL for the folder you want to access.
**Note**: **TLS-encrypted WebDAV URLs use `davs://`** instead of `https://`, so CyVerse URLs should be formatted as:
```sh
davs://data.cyverse.org/dav/iplant/home/
```
4. Click **Connect**.
5. If prompted, enter your **CyVerse username** and **password**.
---
### :material-console: Linux Terminal
To mount a WebDAV folder via the Linux terminal, root or sudo access is required:
1. Install `davfs2` (e.g., on Ubuntu: `sudo apt install davfs2`).
2. Create a mount point:
```sh
mkdir /tmp/data
```
3. Mount the WebDAV directory:
```sh
sudo mount -o gid=,uid= -t davfs /tmp/data
```
- Replace `` with your Linux username.
- Replace `` with the WebDAV folder URL.
4. If prompted, enter your **CyVerse username** and **password**.
---
## :material-account-circle-outline: WebDAV Access Information
Use the following credentials to connect to the Data Store:
| Key | Value |
|-----------------|----------------------|
| `username` | `` |
| `password` | `` |
Use these credentials for anonymous access to the Data Store:
| Key | Value |
|-------------------|-------|
| `username` | `anonymous` |
| `password` | (leave empty) |
---
## :material-web: Data Locations
1. **User Data**
Users can access their data at:
```sh
https://data.cyverse.org/dav/iplant/home//
```
2. **Public Data (Read-Only Access)**
Anonymous users can access public data at:
```sh
https://data.cyverse.org/dav-anon/
```
3. **Community/Project Data**
To access project-specific data stored in iRODS at `/iplant/home/shared//`, use:
```sh
https://data.cyverse.org/dav/iplant/projects//
```
4. **CyVerse Curated Data (DOI-Backed Datasets)**
Access curated datasets with DOIs in the Data Commons at:
```sh
https://data.cyverse.org/dav-anon/iplant/commons/cyverse_curated/
```
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/webdav/file_manager.md){target=_blank} (last source update 2025-03-27), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/data-store/webdav/data-locations/
---
title: "WebDAV data locations"
description: "WebDAV URLs for your own, shared, and public read-only Data Store data at data.cyverse.org."
type: Reference
tags:
- WebDAV
- Data Store
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/webdav/data_location.md"
title: "CyVerse Learning Materials: docs/ds/webdav/data_location.md"
author: "team:cyverse"
last_modified: "2025-03-27T13:28:03-07:00"
---
# WebDAV data locations
When accessing the Data Store via WebDAV, use the following base URLs to locate your data.
## :material-web: User Data
`https://data.cyverse.org/dav/iplant/home/`
- Maps to your iRODS home directory `/iplant/home/`
- Read and write access for the owner
- Not accessible to anonymous users
## :material-web: Public Data (Read-Only)
`https://data.cyverse.org/dav-anon`
- Read-only access
- Available to anonymous users
## :material-web: Community/Project Data
`https://data.cyverse.org/dav/iplant/projects/`
- Maps to the iRODS community-shared data directory: `/iplant/home/shared/`
## :material-web: CyVerse Curated Data (DOI-Backed Datasets)
`https://data.cyverse.org/dav-anon/iplant/commons/cyverse_curated`
- Maps to the iRODS CyVerse curated data directory: `/iplant/home/shared/commons_repo/curated`
- Read-only access
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ds/webdav/data_location.md){target=_blank} (last source update 2025-03-27), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/discovery-environment/overview/
---
title: "Discovery Environment overview"
description: "What the Discovery Environment is, the types of analyses it runs from small interactive jobs to HPC and HTC, and its benefits."
type: Guide
tags:
- Discovery Environment
- Analyses
- Overview
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/index.md"
title: "CyVerse Learning Materials: docs/de/index.md"
author: "team:cyverse"
last_modified: "2025-03-26T10:54:02-07:00"
---
# Discovery Environment overview
## About
Discovery Environment (DE) is an all-purpose data science work bench and fills several major modalities for research computing (Table 1).
The DE can be used for individual research applications that require small to moderate amount of computing in an interactive environment. It can also be used for teaching workshops or classes, where students can each start their own instances with replicated container environments that ensure reproducibility. Last, the DE is connected to the [OpenScienceGrid (High Throughput Computing)](https://osg-htc.org/){target=_blank}, and the [Texas Advanced Computing Center (High Performance Computing)](https://www.tacc.utexas.edu/){target=_blank}, allowing users to launch large HTC and HPC applications with no command line experience.
**Table 1**: Types of research computing environments which the Discovery Environment helps manage for a user.
| Type of Analysis | Cores | RAM | Use Cases |
|------------------|-------|-----|-----------|
| Small | 1-16 | 2-64 GB | Suitable for smaller applications and services, teaching, and long running analyses |
| Moderate | 32-256 | 128 GB - 1 TB | Suitable for larger applications and services that require more resources than a typical laptop |
| High Performance Computing | 1 - 10s thousands | 2 GB to 10s TBs | Used for computationally intensive tasks, such as simulations, data analysis, etc |
| High Throughput Computing | 100s to 100s thousands | Varies | Primarily focused on the efficient execution of a large number of tasks. The emphasis is on high throughput rather than on the speed of any single task |
| Computing Clusters | Varies | Varies | Clusters (Kubernetes) providing Jupyter/Rstudio notebook service to many users |
Discovery Environment's graphical interface is tailored to the needs of researchers who analyze big data but who may also lack command line expertise or the compute resources to run their software tools and analyses at appropriate scale.
## Benefits
- All data in [your own personal space](https://de.cyverse.org/data/ds/){target=_blank}, [Shared with you](https://de.cyverse.org/data/ds/iplant/home?selectedOrder=asc&selectedOrderBy=name&selectedPage=0&selectedRowsPerPage=100){target=_blank}, and public [Community Released](https://de.cyverse.org/data/ds/iplant/home/shared?selectedOrder=asc&selectedOrderBy=name&selectedPage=0&selectedRowsPerPage=100){target=_blank} or [Published (Curated)](https://de.cyverse.org/data/ds/iplant/home/shared/commons_repo/curated?selectedOrder=asc&selectedOrderBy=name&selectedPage=0&selectedRowsPerPage=100){target=_blank} ) stored in CyVerse's cloud-based { width="25" } Data Store and are accessible in the Discovery Environment. See more info about [Data in CyVerse](https://unm-carc.github.io/cyverse/data-store/overview/)
- Internet accessible data can be downloaded / streamed into running analyses.
- All analyses run on CyVerse's Cloud based Kubernetes Cluster enabling you to run analyses that are small (1-16 cores, 2 - 64 GB RAM) to moderate (16-128 cores, 128 GB - 1 TB RAM), and beyond on High Performance Computing and High Throughput Computing resources.
- For most tasks (e.g., launching an app or importing data from a URL) you can log out or navigate to another page or operation after you start the task; an automated email notification is sent to you when the task is completed.
[de]: https://unm-carc.github.io/cyverse/assets/de/logos/deIcon.svg
[home]: https://unm-carc.github.io/cyverse/assets/de/menu_items/homeIcon.svg
[data]: https://unm-carc.github.io/cyverse/assets/de/menu_items/dataIcon.svg
[apps]: https://unm-carc.github.io/cyverse/assets/de/menu_items/appsIcon.svg
[analyses]: https://unm-carc.github.io/cyverse/assets/de/menu_items/analysisIcon.svg
[vice]: https://unm-carc.github.io/cyverse/assets/de/logos/deviceIcon.png
!!! info "Under the hood"
The Sonora web interface, the Terrain API, and the services and Kubernetes cluster that run analyses are described in [Discovery Environment](https://docs.cyverse.org/platform/discovery-environment/){target=_blank} and [System overview](https://docs.cyverse.org/architecture/system-overview/){target=_blank} in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/index.md){target=_blank} (last source update 2025-03-26), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/discovery-environment/login/
---
title: "Logging in to the Discovery Environment"
description: "Sign in to the Discovery Environment and tour its navigation menu: Home, Data, Apps, Analyses, Cloud Shell, Teams, Collections, and Help."
type: Guide
tags:
- Discovery Environment
- Login
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/login.md"
title: "CyVerse Learning Materials: docs/de/login.md"
author: "team:cyverse"
last_modified: "2024-07-16T22:26:28Z"
---
# Logging in to the Discovery Environment
[de]: https://unm-carc.github.io/cyverse/assets/de/logos/deIcon.svg
[data]: https://unm-carc.github.io/cyverse/assets/de/menu_items/dataIcon.svg
[analyses]: https://unm-carc.github.io/cyverse/assets/de/menu_items/analysisIcon.svg
[apps]: https://unm-carc.github.io/cyverse/assets/de/menu_items/appsIcon.svg
[help]: https://unm-carc.github.io/cyverse/assets/de/menu_items/helpIcon.svg
[home]: https://unm-carc.github.io/cyverse/assets/de/menu_items/homeIcon.svg
[profile]: https://unm-carc.github.io/cyverse/assets/de/icons/userIcon.svg
[cloudshell]: https://unm-carc.github.io/cyverse/assets/de/menu_items/webshellIcon.svg
[teams]: https://unm-carc.github.io/cyverse/assets/de/menu_items/teamsIcon.svg
[collections]: https://unm-carc.github.io/cyverse/assets/de/menu_items/bank.svg
When you first open the [![de]{width="25"} Discovery Environment](https://de.cyverse.org){target=_blank}, you'll see the [![home]{width="25"} Home](https://de.cyverse.org/home) Dashboard.
{width="600"}
Discovery Environment Home
The [![home]{width="25"} Home](https://de.cyverse.org/home) Dashboard contains links to [News](https://cyverse.org/news){target=_blank}, recent [YouTube Videos](https://www.youtube.com/c/CyverseOrgProject){target=_blank}, & Featured Apps.
The left side navigation menu shows icons for accessing different parts of the DE. The menu can be expanded by clicking on the three bars in the top left.
{width="200", align=left}
- ![home]{width=20} **Home/Dashboard:** Your main control panel that may display summary widgets, quick links to recent activities, or educational content such as tutorials and webinars.
- ![data]{width=20} **Data:** This interface connects you to the Data Store. Here, you can manage your files, including uploading, downloading, organizing, and sharing data. You'll have access to your personal storage space and shared directories.
- ![apps]{width=20} **Apps:** Discover various applications, including VICE (Visual Interactive Computing Environment) apps for interactive computing sessions. You can browse, search, and launch these applications based on your research needs.
- ![analyses]{width=20} **Analyses:** View and manage your computational tasks. This section logs your history of analysis jobs, allowing you to monitor current processes, review completed ones, and access resulting data.
- ![cloudshell]{width=20} **Cloud Shell:** Access a Linux shell environment directly within the DE. This feature enables advanced users to perform command-line operations without leaving the platform.
- ![teams]{width=20} **Teams:** Create and manage collaboration groups. Teams allow you to group together with other users for easier sharing of data, analyses, and other collaborative efforts.
- ![collections]{width=20} **Collections:** Explore public collections of data and apps curated by other users or the CyVerse team. This resource can be invaluable for finding information relevant to your studies.
- ![help]{width=20} **Help:** Access various support materials, including FAQs, guides, and contact information for direct assistance from the CyVerse support team.
Sign in from the upper right corner of the DE and click the ![profile]{width="25"} profile icon or clicking the [Login](https://de.cyverse.org/#){target=_blank} link. If you attempt to view the Data Store or launch an App, you will see a pop-up:
{width="300"}
When signing in you will be redirected to our Authentication Service. Enter your CyVerse username and password.
If you don't have an account yet or you've forgotten your password, you can visit {target=_blank} to create an account.
{width="600"}
After logging in, you'll be returned to the [![home]{width="25"} Home](https://de.cyverse.org/home) Dashboard.
If you were already on the [![apps]{width="25"} Apps](https://de.cyverse.org/apps){target=_blank} or [![data]{width="25"} Data](https://de.cyverse.org/data){target=_blank} when you logged in, you'll return to that page.
You can take a short tour of the DE's main features by clicking the [help icon ![help]{width="25"}](https://de.cyverse.org/help){target=_blank} in the left sidebar and selecting "Product Tour".
!!! info "Under the hood"
Sign-in is handled by CyVerse's Keycloak service, which can broker institutional (CILogon) and OAuth logins; see [Authentication](https://docs.cyverse.org/platform/authentication/){target=_blank} in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/login.md){target=_blank} (last source update 2024-07-16), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/discovery-environment/using-apps/
---
title: "Using apps in the Discovery Environment"
description: "Browse, sort, filter, and inspect apps in the Discovery Environment, and understand VICE and HPC apps and HPC queues."
type: Guide
tags:
- Discovery Environment
- Apps
- HPC
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/using_apps.md"
title: "CyVerse Learning Materials: docs/de/using_apps.md"
author: "team:cyverse"
last_modified: "2026-03-16T13:18:13-07:00"
---
# Using apps in the Discovery Environment
[de]: https://unm-carc.github.io/cyverse/assets/de/logos/deIcon.svg
[apps]: https://unm-carc.github.io/cyverse/assets/de/menu_items/appsIcon.svg
[vice]: https://unm-carc.github.io/cyverse/assets/de/logos/deviceIcon.svg
You can select from several hundred applications (Apps) available in the [![de]{width="25"} Discovery Environment](https://de.cyverse.org){target=_blank} when you are ready to analyze your data.
!!! tip "When launching [![apps]{width="10"} Apps](https://de.cyverse.org/apps){target=_blank}, you can log out or navigate to another page or operation after you start the task; an automated email notification is sent to you when those tasks are completed"
## Browsing Apps in the Discovery Environment
You must be logged in to browse and use apps.
1. Click in the left sidebar of the DE to see the [![apps]{width="25"} Apps](https://de.cyverse.org/apps){target=_blank} view. When you first access the Apps view, you may be prompted to log in. After logging in, you will see a screen that looks something like this:
{ width="600" }
The Apps page.
## Sorting and Filtering Apps in the Discovery Environment
To sort the list of apps in ascending or descending order by app name, the name of the person who integrated the app, or its average rating, click on the column headings.
You can navigate between pages and change how many apps are listed on a page by using the < or > controls at the bottom of the page.
By default, the Apps view displays "**Featured Apps**" which are interactive.
All "**Public Apps**" are available to you.
With hundreds of apps and sometimes many versions of an app in the DE, you may want to view a subset of all available apps. There are two ways to do this.
First, in the upper left corner of the [![apps]{width="25"} Apps](https://de.cyverse.org/apps){target=_blank} view, the currently active subset of apps is shown as the primary filter.
Click the drop-down arrow next to the currently active subset to select a different apps subset to display:
The currently selected app subset is highlighted in gray. The available app subsets are:
| Application type | Description |
|------------------|-------------|
| Apps under development | Apps that you have added to the DE that have not been made public |
| Favorite apps | Apps that you have marked as favorite apps in the DE |
| My public apps | Apps that you have added to the DE that have been made publicly available |
| Shared with me | Apps that other users have shared with you |
| High-Performance Computing | Apps that run at the Texas Advanced Computing Center using the Tapis API |
| Browse All Apps | All apps available to you in the DE |
You can further reduce the list of the apps displayed by selecting a filter.
Click the drop-down arrow in the Filter control (upper right corner of the Apps view) to select the type of apps you'd like to see in the listing:
The currently selected filter is displayed in the Filter control itself.
If no filter is selected, the control will be empty. The currently available app filters are:
| Application filter | Description |
|--------------------|-------------|
| HPC | High-Performance Computing apps that run using the Tapis API |
| DE | Executable (non-interactive apps) that run on CyVerse computing resources |
| VICE | Interactive development environments (e.g., Jupyter, RStudio, R Shiny) and other apps with their own interactive interfaces |
| Open Science Grid (OSG) | Executable (non-interactive apps) that run on OSG resources |
The app filter you selected will be displayed in the Filter control.
## Viewing App Details in the Discovery Environment
When you've found an app of interest, select it by clicking the checkbox to the left of the app name.
A *Details* button will appear in the upper right corner of the Apps view, just to the right of the Filter control.
Click the Details button to see additional information about the app (e.g., description, number of times run, etc.).
The Details panel has several controls available.
Click the Heart icon to add that app to your list of favorite apps (to remove from your favorite list, click the heart again).
The heart will be solid blue if the app is already on your list of favorites.
Click the Link icon to display a link to the app that you can copy and share with other CyVerse users.
The Stars icon labeled `Your rating` allows you to rate the app.
The `Tools used by this App` tab contains information about the underlying tools (steps) the app uses to perform an analysis.
To dismiss the App Details view, click anywhere outside the panel.
??? tip "Create a Favorites list"
Favorite your frequently used apps to make them easier and faster to find next time.
## About VICE Apps in the Discovery Environment
One type of app that you can filter for in the [![de]{width="25"} Discovery Environment](https://de.cyverse.org){target=_blank} are [![vice]{width="25"}](https://de.cyverse.org){target=_blank} (VICE stands for Visual Interactive Computing Environment). VICE apps are interactive apps that include a Graphical User Interface (GUI) or an Integrated Development Environment (IDE) such as Project Jupyter, RStudio, or remote desktops to the DE.
You must request special access and be approved to use VICE apps through the CyVerse User Portal .
## About HPC Apps in the Discovery Environment
Most DE apps that are listed in the High-Performance Computing (HPC) category,
as well as CyVerse apps which run through the [Tapis API](https://tapis.readthedocs.io){target=_blank},
run at [TACC](https://tacc.utexas.edu){target=_blank} (the Texas Advanced Computing Center),
and mainly on their [Stampede3 system](https://docs.tacc.utexas.edu/hpc/stampede3/){target=_blank}.
Access to this powerful resource is made available through a grant from the National Science Foundation.
Stampede3 allocation requests must be made through the NSF's [ACCESS](https://allocations.access-ci.org){target=_blank} project.
!!! tip "You must log in to the Tapis API server in order to view and use the list of HPC apps."
If you have not yet authenticated with Tapis
(you'll only have to do it once, or after resetting your HPC Token under `Settings`),
select the "High-Performance Computing" category in the Apps listing page,
then log in with your CyVerse credentials when prompted to authenticate with the cyverse.tapis.io server.
### Authenticating with TAPIS and Stampede3
In order to use HPC Apps in the DE that run on TACC's Stampede3 system,
follow their [Getting Started guide](https://tacc.utexas.edu/use-tacc/getting-started/){target=_blank}
to register for a TACC account and to request a Stampede3 allocation.
Be sure to choose the same username for your TACC account as your CyVerse username.
After your TACC account is activated, navigate to [cyverse.tapis.io](https://cyverse.tapis.io/){target=_blank}
and log in at that page with your **CyVerse credentials**
(disregard the help messages in the login form that asks for your TAPIS name and password).
Then navigate to the Systems page and find the public
[stampede3 system](https://cyverse.tapis.io/#/systems/stampede3){target=_blank}.
Near the top of the `stampede3` details page,
there will be a display that checks if you have "Authenticated"
with this `stampede3` system with your **TACC credentials**.
If this section displays a message that you are unauthenticated,
then it will provide an "Authenticate" link that you can select
which will display a form for you to enter your **TACC account password**.
{ width="338" }
Unsuccessful Stampede3 Authentication Check
After entering your TACC account password,
refreshing this page should display a successful authentication check.
{ width="338" }
Successful Stampede3 Authentication Check
This is a one-time authentication required for DE HPC apps running on the `stampede3` system.
### Understanding HPC queues
In order to fairly distribute this high-demand resource,
TACC follows allocation policies that limit how long any single analysis can be run
(usually 24 or 48 hours, depending on the queue),
how many analyses a user can have running simultaneously,
and the total amount of computational time any one user can access over the course of a year.
Analyses (also known as jobs) submitted through the CyVerse DE run at TACC using the same queues as every other scientist in the country uses.
Thus, if there are many analyses or a few very large analyses in the queue,
the wait time for each analysis can be very long,
up to several days for certain apps.
Queues on HPC systems are much like queues at the coffee shop:
the first analysis submitted is the first one to run.
However, to efficiently exploit resources,
HPC queues also have features similar to amusement parks that squeeze single riders in with larger groups.
On an HPC system, this consists of scheduling jobs that are shorter or that use fewer nodes into smaller blocks that can be placed in between longer jobs.
Supercomputing centers generally have more than one supercomputer,
and the supercomputers have multiple queues for different types of analyses
(e.g., serial, parallel, large memory).
Each center/computer/queue has its own rules and algorithms for ensuring fair and efficient allocation of resources.
Users or groups who have very large computational needs are likely to run into bottlenecks using standard CyVerse infrastructure.
We recommend that these users
[apply for their own ACCESS allocation](https://allocations.access-ci.org){target=_blank},
which will allow them to run CyVerse tools and applications at TACC with fewer restrictions.
Users or groups with very large computational needs should first apply for a startup allocation and use it to benchmark their jobs,
thereby collecting data on efficiency of resource use which must be part of a full ACCESS allocation request.
Want to learn more about ACCESS?
Visit the [ACCESS home page](https://access-ci.org){target=_blank}.
## Advanced Features in the Discovery Environment
The Discovery Environment also supports advanced features for apps such as integrating different types of apps into the DE, creating and running containers, and using Application Programming Interfaces (APIs) for programmatic backend access to CyVerse services.
For how-to information on these features, see our [Developer Manuals](https://unm-carc.github.io/cyverse/developers/manuals/), [Extending VICE Apps](https://unm-carc.github.io/cyverse/discovery-environment/vice/extend-apps/), and our [Powered By](https://unm-carc.github.io/cyverse/getting-started/powered-by/) documentation.
!!! info "Under the hood"
The app catalog, the executable, interactive (VICE), and HPC job types, and the services that launch them are described in [Discovery Environment](https://docs.cyverse.org/platform/discovery-environment/){target=_blank} and [System overview](https://docs.cyverse.org/architecture/system-overview/){target=_blank} in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/using_apps.md){target=_blank} (last source update 2026-03-16), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/discovery-environment/managing-analyses/
---
title: "Managing analyses in the Discovery Environment"
description: "Check the status of, relaunch, inspect, share, and cancel analyses in the Discovery Environment, and use CPUs efficiently."
type: Guide
tags:
- Discovery Environment
- Analyses
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/managing_analyses.md"
title: "CyVerse Learning Materials: docs/de/managing_analyses.md"
author: "team:cyverse"
last_modified: "2023-07-28T01:59:34+05:30"
---
# Managing analyses in the Discovery Environment
An analysis is the product of a launched app that has completed its computation of input data. The Discovery Environment maintains a history
of all your analyses, including a unique analysis ID, launch date, input files, and other details.
## Browsing and Checking the Status of Analyses in the Discovery Environment
1. Open the **Analyses** view by clicking on the (analyses icon) on the left sidebar of the DE interface to monitor the status of your submitted
analysis. The analysis launched most recently will be at the top of the list.
2. Analyses can be sorted by Name, Start date, End date or Status. To sort your analyses, hover your cursor over the name of the column you
wish to sort by and click on the arrow that appears beside the column name.
??? tip "Analyses Status"
In the DE, an analyses can have one of the following status messages:
- **Submitted**: The analysis has been queued for running on CyVerse resources.
- **Running**: The analyses is now running - for most apps (non-interactive) the analysis will run until it is completed or a time limit is reached. For interactive (VICE) applications, you may now access your interactive application (check your notification icon for a link or click the link-out icon ).
- **Completed**: The application is completed and any logs and results have been written to the data store. Access the outputs by clicking the folder icon )
- **Canceled**: The analyses has been cancelled.
- **Failed**: The analyses has resulted in an error.
3. To filter your analyses by user, click on the View dropdown menu in the upper left corner and select either 'My analyses' or 'Shared with me'. The default view is 'My analyses'.
4. To further filter your analyses by app type, click on the Filter dropdown menu and select the type of analyses you would like to see (i.e., HPC, DE, VICE, or OSG).
5. To open and view the output folder of a completed analysis, click on the output folder icon at the right side of that particular analysis.
## Relaunching an Analyses in the Discovery Environment
You can relaunch an analyses to load an analyses you have previously done. Your analyses will load with the same inputs and parameters as previously used and you will then have the option to some, all, or none of the of those settings.
*To relaunch*
1. Select an analysis from your history.
2. Click the relaunch icon ().
3. Alter any desired parameters and launch the application.
## Viewing Analyses Details in the Discovery Environment
Click the "Details" button or the (info icon) to view details of the analysis (e.g. parameters used).
## Sharing an Analyses in the Discovery Environment
Clicking the "Share" or "Add To Bag" button to share an analysis and its results with another CyVerse user.
## Additional Analyses Actions in the Discovery Environment
Clicking the "More Actions" button allows you to perform the following actions:
- **Go to Output folder**: View the logs and outputs of a completed analysis
- **Relaunch**: Relaunch an analysis (with the option to edit parameters)
- **Rename**: Rename an analysis
- **Update Comments...**: Add or edit comment notation on an analysis
- **Go to analyses**: View an interactive analysis in a new tab
- **Extend Time Limit**: Extend the time limit of an interactive analysis
- **Terminate**: Stop a submitted or running analysis
- **Delete**: Delete an analysis from your history
- **Add to Bag**: Add to a "bag" for sharing
## Using CPUs Efficiently: Best Practices
There are two ways to reduce the number of CPU hours that are consumed.
1. **Request Fewer CPUs during Analysis Submission**
- In the "Advanced Settings" tab of the analysis launch wizard, you can modify the "Maximum CPU Cores" setting.
- Selecting 0 will automatically select the default setting, which is currently set to 4 CPUs.
- Selecting 1 will request only 1 CPU. For multithreaded applications, it's advisable to select more than 1 CPU, depending on the specific application's requirements.
2. **Terminate Completed Analyses Promptly**
- As soon as an analysis is complete, terminate it by clicking the red X button next to the analysis in the analysis listing.

- Note: Leaving an analysis running while it's not doing anything is one of the quickest ways to use up CPU hours.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/managing_analyses.md){target=_blank} (last source update 2023-07-28), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/discovery-environment/bags/
---
title: "Sharing and using bags in the Discovery Environment"
description: "Collect data, apps, and analyses in a bag to share them with collaborators or download them together from the Discovery Environment."
type: Guide
tags:
- Discovery Environment
- Sharing
- Bags
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/bags.md"
title: "CyVerse Learning Materials: docs/de/bags.md"
author: "team:cyverse"
last_modified: "2022-03-21T16:50:12-05:00"
---
# Sharing and using bags in the Discovery Environment
[bag]: https://unm-carc.github.io/cyverse/assets/de/icons/bagIcon.svg
[de]: https://unm-carc.github.io/cyverse/assets/de/logos/deIcon.svg
[data]: https://unm-carc.github.io/cyverse/assets/de/menu_items/dataIcon.svg
[analyses]: https://unm-carc.github.io/cyverse/assets/de/menu_items/analysisIcon.svg
[apps]: https://unm-carc.github.io/cyverse/assets/de/menu_items/appsIcon.svg
[help]: https://unm-carc.github.io/cyverse/assets/de/menu_items/helpIcon.svg
[home]: https://unm-carc.github.io/cyverse/assets/de/menu_items/homeIcon.svg
[profile]: https://unm-carc.github.io/cyverse/assets/de/icons/userIcon.svg
You can share data, apps, and analyses using the sharing features with one or more CyVerse users through the [![de]{width="25"} Discovery Environment](https://de.cyverse.org){target=_blank}.
The ![bag]{width="25"} Bag is a handy feature in the DE that you can use to gather and download or share multiple data files, apps, and analyses or any combination of those resources with another user(s).
There is no limit to the number of items you can put in a bag but you can only share files you own.
## Sharing Data in the Discovery Environment
You must be logged in to share resources.
1. Open the [![data]{width="25"} Data](https://de.cyverse.org/data){target=_blank} view of your home directory by clicking on in the left sidebar.
2. Select the data resource(s) you wish to share; then click the "Share" button.
3. In the Sharing dialog that opens, check that the resources you wish to share are shown.
4. In the Search checkbox, search for the CyVerse user(s) you wish to share with by typing their full CyVerse username or email address.
Subsequent searches will more quickly find names from your previous "Shared with" lists.
5. Next, under "Permission", choose which type of permission to grant the person(s) you are sharing resources with. You can also "Remove" access using the Permission dialog box.
6. When you are finished, click "Done" to begin sharing. The user(s) will be notified that resources have been shared with them and will see the shared item(s) in their "Shared With Me" folder when they log in.
??? tip "Managing Data Permissions"
| Permission level | Read | Download/Save | Metadata | Rename | Move | Delete |
|:----------------:|:----:|:-------------:|:--------:|:------:|:----:|:------:|
| Read | **X** | **X** | **View** | | | |
| Write | **X** | **X** | **Add/Edit** | | | |
| Own | **X** | **X** | **Add/Edit** | **X** | **X** | **X** |
## Sharing Apps in the Discovery Environment
You must be logged in to share resources.
1. Open the [![apps]{width="25"} Apps](https://de.cyverse.org/apps){target=_blank} view by clicking on the (app icon) in the left sidebar.
2. Select your app(s) or app(s) you are building that you wish to share with another user(s) or your team; then click the Share button.
3. In the Sharing dialog that opens, check that the app(s) you wish to share is shown.
4. In the Search box at the top of the page, start typing the CyVerse username, team name, or email address of the CyVerse user(s) with whom you want to share. Search will start when you enter at least three characters. There is no limit to how many users you can share files and analyses with.
5. Next, under "Permission", choose which type of permission you want to grant the person(s) or team you are sharing the app(s) with.
6. Once you are finished, click "Done" to begin sharing. The user(s) will be notified that app(s) have been shared with them when they log in.
??? tip "Managing App Permissions"
Permissions (based on UNIX permissions) are described in this chart:
| Permission level | Launch | Edit | Share | Make Public |
|:----------------:|:------:|:----:|:-----:|:-----------:|
| Read | **X** | | | |
| Write | **X** | **X** | | |
| Own | **X** | **X** | **X** | **X** |
## Sharing Analyses in the Discovery Environment
You must be logged in to share resources.
1. Open the [![analyses]{width="25"} Analyses](https://de.cyverse.org/analyses){target=_blank} view by clicking on the (analyses icon) in the left sidebar.
2. Select one or more analyses you wish to share with another user(s); then click the Share button in the upper right corner of the page.
3. In the Sharing dialog that opens, ensure that the analyses you wish to share are shown under Resources.
4. In the Search box at the top of the page, search for the CyVerse user(s) you wish to share with by typing their full CyVerse username, team name or email address. Search will begin when you have typed at least three characters. Click the desired user(s). There is no limit to how many users you can share analyses with.
5. Next, under "Permission", choose which type of permission to grant the person(s) you are sharing the analyses with.
6. When you are finished, click "Done" to begin sharing. The user(s) will be notified that analyses have been shared with them when they
log in.
??? tip "Manage Analyses Permissions"
| Permission level | Read | Download/Save | Metadata | Rename | Move | Delete |
|:----------------:|:----:|:-------------:|:--------:|:------:|:----:|:------:|
| Read | **X** | **X** | **View** | | | |
| Write | **X** | **X** | **Add/Edit** | | | |
| Own | **X** | **X** | **Add/Edit** | **X** | **X** | **X** |
## Using a Bag to Share or Download in the Discovery Environment
You can share or download multiple items using the ![bag]{width="25"} feature in the Discovery Environment.
You must be logged in to use a bag. There is no limit to the number of items you can share in a bag.
1. To share file(s) using a ![bag]{width="25"} bag, open the Data, Apps or the Analyses view (or each consecutively), select one or more files, and click "Add to Bag" in the upper right corner. A red dot will appear on the (bag icon) to show how many resources are currently in the bag.
2. When you've finished adding all the files (data, apps or analyses) you want to share in the bag, share the bag with another CyVerse
user(s) by clicking on the (bag icon). In the dialog box that opens, all the files you have put in the bag are listed by default. Use the
dropdown arrow to show downloadable or shareable files.
3. You can "Share" the contents of the bag by clicking on the "Share" button. Another dialog box will open where you can set Permissions
for the user(s) with whom you are sharing files.
4. To download the bag's contents to your computer, click "Download" and then click on each of the links for the files.
5. Sharing or downloading the contents of a bag does not empty the bag. You can share the same contents with another user(s); to empty the
bag, click on the (bag icon) then click the "Clear" button.
??? tip "Emptying Bags"
There is no prompt or warning once you click "Clear", so the bag will be emptied immediately.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/bags.md){target=_blank} (last source update 2022-03-21), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/discovery-environment/create-apps/
---
title: "Create your own apps in the Discovery Environment"
description: "Integrate a containerized tool and build an app interface for it in the Discovery Environment, then test, share, and publish it."
type: Guide
tags:
- Discovery Environment
- App Integration
- Docker
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/create_apps.md"
title: "CyVerse Learning Materials: docs/de/create_apps.md"
author: "team:cyverse"
last_modified: "2026-03-16T13:18:13-07:00"
---
# Create your own apps in the Discovery Environment
## Why use the DE?
- Use hundreds of bioinformatics Apps without the command line (or
with, if you prefer)
- Batch and interactive modes
- Seamlessly integrated with data and high performance computing --
not dependent on your hardware
- Create and publish Apps and workflows so anyone can use them
- Analysis history and provenance -- "avoid forensic bioinformatics"
- Securely and easily manage, share, and publish data
## :material-grid: Apps vs :octicons-container-24: Tools
**:octicons-container-24: Tool:** A "tool" is essentially a container image. The image must be hosted on a public container registry (like the Docker Hub, Biocontainers, NVIDIA GPU Cloud, etc.) or the [CyVerse Harbor Registry](https://harbor.cyverse.org){target=_blank}. Each "Tool" uses a template editor which is internal to the DE for defining the `registry/image-name:tag` of the image, its working directory, and runtime environment.
**:material-grid: App:** a simple graphic interface for running the Tool in the DE with any special commands or input requirements defined in the App template. Apps can have different flavors:
- **Executable**: user starts an analysis and when the analysis
finishes they can find the output files in their `/analyses`
folder
- **DE**: run locally on our cluster
- **HPC**: run at the Texas Advanced Computing Center (TACC)
- **OSG**: run on the Open Science Grid
- **Interactive**: also called Visual and Interactive Computing
Environment (VICE). Allows users to open Integrated Development
Environments (IDEs) including RStudio, Project Jupyter and RShiny
and work interactively within them.
![][hpc]{width=20} **HPC apps** must be [created through the Tapis API](https://unm-carc.github.io/cyverse/discovery-environment/create-hpc-apps/).
![][apps]{width=20} For all other app types,
**the ([containerized](https://cyverse-de-manual.readthedocs-hosted.com/en/latest/new-appInterfacechildpages/DockerizingTools.html){target=_blank}) tool must be [integrated into the CyVerse DE first](#adding-a-tool). Then an [app (interface) can be built](#building-an-app-for-your-tool) for that tool.**
## Adding a Tool
!!! tip "Check for existing Tools first"
It is a good idea to check if the tool you want is already integrated before you start. The tool may be there already and you can build an app using it.
1. Log into the [![][de]{width=25}](https://de.cyverse.org){target=_blank} [Discovery Environment](https://de.cyverse.org){target=_blank}.
2. Click the [![][apps]{width=20}](https://de.cyverse.org/apps){target=_blank} [Apps](https://de.cyverse.org/apps){target=_blank} and click on the "Manage Tools" wrench icon.
3. You'll see a list of all of the tools in the DE. You can search for the tool you want to see if it has already been integrated. If you can't find it, click on "More Actions"
{width="100"}
and select "Add Tool".
{width="100"}
The template editor will open with the following fields will need to be completed (note: not all fields are required) before the Tool can be run.
{width="600"}
**Add Tool**
- `Tool name` is the name of the tool. This will appear in the DE's tool listing dialog. This is mandatory field.
- `description` is a brief description of the tool. This will appear in the DE's tool listing dialog.
- `version` is the version of the tool. This will appear in the DE's tool listing dialog. This is mandatory field.
- `Type` is the type of tool. For command line applications, choose "executable".
**Container Image**
- `Image name` is the name of the image and its public registry. This is mandatory field.
- `Tag` is the image tag. If you don't specify the tag, the DE will look for the `latest` tag which is the default tag.
- `Docker Hub URL` is the URL of the image on Dockerhub.
- `Entrypoint` is the Entrypoint for your tool. Entrypoint should be present in the Docker image, and if not, you should specify it here.
- `Working Directory` is the working directory of the tool. `WORKDIR` should be present in the Docker image, and if not, you should specify it here.
**Restrictions**
- `Max CPU Cores` is the number of cores for your tool, e.g., 16
- `Memory Limit` is the memory for your tool, e.g., 64 GB
- `Min Disk Space` is the minimum disk space for your tool, e.g., 200 GB
## Building an App for Your Tool
You can build an app for any tool that:
- is private to you
- is shared with you
- is public
Step 1: App Info
From the 'Apps' tab click on the 'Create' button in the top right corner and select 'Add App'. Choose an informative app name and description (eg. tool
name and version). Select the tool you want to build the app on buy clicking the 'select' button. This will open the 'search tools' window. Search for and select your tool.
{width="600"}
{width="600"}
Step 2: Parameters
Divide the app into sections appropriate for that tool (input, options and output are usually
sufficient sections for simple apps). You can add a section by clicking on the 'Add Section'. Once you have added a section you can edit the name by clicking on the pencil icon (right side). Within a section you can add the parameters necessary for your tool by clicking on 'Add Parameter' and choosing the type of parameter you want to add (e.g. input file). For each option you add, you will need to specify what the option is,
the argument option (if there is one) and whether that option is required. If an
option is not required be sure to check the 'exclude if nothing is
entered' box. For tools that have positional arguments (no argument option, eg.
-i) you can leave argument option blank but you will need to make sure your arguments are in the proper order in step 4.
{width="600"}
**Note:**
Although it is best to add all of the options for your tool, as it makes
the app the most useful, you can expose as many or as few options as you
like (as long as you add all the required options).
Step 3: Preview App
Make sure your app looks the way you want it to and that you have included all of the required options. If you need to make changes use the back button to return to the previous step.
{width="600"}
Step 4: Command Line Order
This will provide a preview of what your options will look like on the command line. In the list of options below, use the up and down arrows to the right of the option to move it up or down in the list. You should see these changes reflected in the command line preview box. This order is especially important if your tool uses positional arguments.
{width="600"}
Step 5: Completion
Click 'Save' (upper right) to save your work. Then click 'Launch App' at the bottom of the page and test your app with appropriate data.
If you need to make changes to your app after testing, you can find it under the 'Apps under development' section of the 'Apps' tab. Click on the three dots menu (to the right of your app) and select 'edit app'. This will re-open the apps editor and allow you to make changes.
{width="600"}
Once you know your app works correctly you can share or publish it as
you wish. Public apps must have example data located in an appropriately
named folder here:
`/iplant/home/shared/iplantcollaborative/example_data`
All public apps also have a brief documentation page on the [CyVerse
Wiki](https://cyverse.atlassian.net/wiki/spaces/DEapps/pages/241882146/List+of+Applications)
To publish your app click on the three dots menu (at the right of your app)
and select 'Publish'. You will need to supply:
- location of the example data
- brief description of inputs, required options and outputs
- link to CyVerse Wiki documentation page
- link to documentation for the tool (provided by the developers)
[de]: https://unm-carc.github.io/cyverse/assets/de/logos/deIcon.svg
[apps]: https://unm-carc.github.io/cyverse/assets/de/menu_items/appsIcon.svg
[hpc]: https://unm-carc.github.io/cyverse/assets/de/icons/HPCIcon.svg
!!! info "Under the hood"
Tools and apps are stored and launched through the Terrain API. To script app integration or contribute to the platform, see [Terrain](https://docs.cyverse.org/api/terrain/){target=_blank}, [Endpoint index](https://docs.cyverse.org/api/endpoint-index/){target=_blank} and [Developer guide](https://docs.cyverse.org/development/developer-guide/){target=_blank} in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/create_apps.md){target=_blank} (last source update 2026-03-16), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/discovery-environment/create-hpc-apps/
---
title: "Create HPC apps for use in the Discovery Environment"
description: "Define, register, and share Tapis v3 apps that run at TACC so they appear as HPC apps in the Discovery Environment."
type: Guide
tags:
- HPC
- Tapis
- App Integration
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/create_hpc_apps.md"
title: "CyVerse Learning Materials: docs/de/create_hpc_apps.md"
author: "team:cyverse"
last_modified: "2026-03-16T13:18:13-07:00"
---
# Create HPC apps for use in the Discovery Environment
In order to create a High-Performance Computing (HPC) app for the DE that runs at
TACC (the [Texas Advanced Computing Center](https://tacc.utexas.edu){target=_blank}),
the app must be created through the
[Tapis v3 API](https://cyverse.tapis.io){target=_blank}.
The [Tapis v3 App docs](https://tapis.readthedocs.io/en/latest/technical/apps.html){target=_blank}
can be referenced along with this guide when creating a new HPC app for the DE,
and for other details beyond the scope of this guide.
If you have created HPC apps with Tapis v2 (formerly known as Agave) for the DE in the past, then the guide
[Tapis v2 (Agave) App migration to v3 for the Discovery Environment](https://docs.cyverse.org/api/tapis-v2-v3-migration/){target=_blank}
can help you migrate your v2 apps so they can be used in the DE once again.
## Tapis `curl` command setup
Please note that the API server at [cyverse.tapis.io](https://cyverse.tapis.io/){target=_blank}
does provide its own UI for creating apps,
but it is limited in scope and does not allow for adding JSON `notes` fields,
which are used by the DE to display different parameter field types beyond simple text field inputs
(such as checkboxes, selection lists, and file/folder inputs).
It also does not currently allow editing of existing apps.
So this guide will use `curl` commands to interface with the `cyverse.tapis.io` API
(but there are other options detailed in the
[Tapis "Getting Started" guide](https://tapis.readthedocs.io/en/latest/getting-started/){target=_blank}).
The following shell environment variables and `curl` command aliases
will be used in the examples throughout the rest of this guide.
### Set env var `TAPIS_ACCESS_TOKEN`
Assuming the env vars `CYVERSE_USERNAME` and `CYVERSE_PASSWORD` are set,
this command will call the
[/v3/oauth2/tokens endpoint](https://tapis-project.github.io/live-docs/?service=Authenticator#tag/Tokens/operation/create_token){target=_blank}
and store the resulting Tapis JWT in the env var `TAPIS_ACCESS_TOKEN`:
export TAPIS_ACCESS_TOKEN=$(curl -H "Content-Type: application/json" -s -d "{\"username\": \"$CYVERSE_USERNAME\", \"password\": \"$CYVERSE_PASSWORD\", \"grant_type\": \"password\" }" https://cyverse.tapis.io/v3/oauth2/tokens | jq -r .result.access_token.access_token)
Note the use of the [jq](https://jqlang.github.io/jq/){target=_blank} command to parse out the
`access_token` from the API response.
### Set a `curl` alias that references the `TAPIS_ACCESS_TOKEN`
alias curl-tapis='curl -H "X-Tapis-Token: $TAPIS_ACCESS_TOKEN" -H "Content-Type: application/json"'
### List Systems
If the env vars and alias above are set up correctly,
then a command like the following
(which calls the [/v3/systems endpoint](https://tapis-project.github.io/live-docs/?service=Systems#tag/Systems/operation/getSystems){target=_blank})
should return some results.
curl-tapis -s "https://cyverse.tapis.io/v3/systems?select=id,enabled,systemType,owner,canExec,isPublic&listType=ALL"
Note that these API endpoints return JSON output,
which can be formatted by another command like [jq](https://jqlang.github.io/jq/){target=_blank},
but the remaining example commands in this guide will omit piping outputs to formatting commands.
## Tapis v3 App Definition
First consider a Tapis v3 app definition like the following:
```json
{
"id": "CyVerse-QA-Test-App",
"version": "0.1",
"tenant": "cyverse",
"description": "QA Test app for all parameter types.",
"runtime": "ZIP",
"containerImage": "tapis://data.cyverse.org/home/api/v3/apps/CyVerse-QA-Test-App-0.1.tgz",
"jobType": "BATCH",
"tags": ["QA", "test"],
"notes": {
"name": "CyVerse QA Test App",
"helpURI": "https://learning.cyverse.org/de/create_hpc_apps/",
"outputs": [
{
"name": "requiredOutput",
"description": "Output File command line option.",
"arg": "out.txt",
"details": {
"label": "Output File"
}
}
]
},
"jobAttributes": {
"execSystemId": "stampede3",
"maxMinutes": 100,
"parameterSet": {
"envVariables": [
{
"key": "NEW_ENV_VAR",
"value": "testing env",
"inputMode": "INCLUDE_ON_DEMAND"
}
],
"appArgs": [
{
"name": "requiredOutput",
"description": "Output File command line option.",
"arg": "out.txt",
"inputMode": "REQUIRED",
"notes": {
"details": {
"label": "Output File"
},
"value": {
"default": "out.txt"
}
}
},
{
"name": "textInput",
"description": "Single-line text.",
"arg": "default_value",
"notes": {
"details": {
"label": "Text Input"
}
}
},
{
"name": "flagInput",
"description": "Checkbox input.",
"arg": "-f",
"notes": {
"details": {
"label": "Flag Input"
},
"value": {
"type": "flag"
}
}
},
{
"name": "flagInputAlt",
"description": "Another Checkbox input.",
"arg": "-f2",
"notes": {
"details": {
"label": "Flag Input 2"
},
"value": {
"type": "string"
},
"semantics": {
"ontology": ["xs:boolean"]
}
}
},
{
"name": "integerInput",
"description": "Integer input.",
"notes": {
"details": {
"label": "Integer"
},
"value": {
"type": "number"
},
"semantics": {
"ontology": ["xs:int"]
}
}
},
{
"name": "doubleInput",
"description": "Decimal input.",
"notes": {
"details": {
"label": "Decimal"
},
"value": {
"type": "number"
}
}
},
{
"name": "listInput",
"description": "List input.",
"notes": {
"value": {
"visible": true,
"required": false,
"type": "enumeration",
"order": 0,
"default": null,
"enum_values": [
{
"--list val1": "Value 1"
},
{
"--list val2": "Value 2"
},
{
"--list val3": "Value 3"
}
]
},
"details": {
"label": "Selection List"
}
}
},
{
"name": "requiredInput",
"description": "Required input file.",
"inputMode": "REQUIRED",
"notes": {
"details": {
"label": "Required Input File"
}
}
},
{
"name": "optionalInput",
"description": "Not required, excluded if empty.",
"notes": {
"details": {
"label": "Optional Input File"
}
}
}
]
},
"fileInputs": [
{
"name": "requiredInput",
"description": "Required input.",
"inputMode": "REQUIRED",
"targetPath": "*"
},
{
"name": "optionalInput",
"description": "Not required, excluded if empty.",
"targetPath": "*"
},
{
"name": "fixedInput",
"description": "Fixed input.",
"inputMode": "FIXED",
"sourceUrl": "tapis://data.cyverse.org/home/shared/cyverse_training/example/coffee_cake.txt",
"targetPath": "*"
}
]
}
}
```
If this app JSON was saved in a file named `app.json`,
then this app can be created with a command like the following,
which calls the
[POST /v3/apps endpoint](https://tapis-project.github.io/live-docs/?service=Apps#tag/Applications/operation/createAppVersion){target=_blank}:
curl-tapis -s -d @app.json https://cyverse.tapis.io/v3/apps
If the app needs to be updated,
save the updates to the same file and update the app with the
[PUT /v3/apps endpoint](https://tapis-project.github.io/live-docs/?service=Apps#tag/Applications/operation/putApp){target=_blank}:
curl-tapis -s -X PUT -d @app.json "https://cyverse.tapis.io/v3/apps/CyVerse-QA-Test-App/0.1"
## Tapis v3 App Details
### The `notes` Field
The DE will use an app's `notes` fields to decide how to display and validate certain app fields and parameters,
such as app and parameter `name` and `label` display fields,
if those `notes` fields are formatted
[similar to Tapis v2 app fields](https://docs.cyverse.org/api/tapis-v2-v3-migration/){target=_blank}.
The DE will use the app's `notes.label` or `notes.name`,
followed by the app's `version`,
when displaying the app name in listings or the in the app launch form.
If no `label` or `name` is found in the app's `notes`,
then the DE will simply use the app's `id`.
### Parameters and App Args
Now take the following `appArgs` parameter, for example:
```json
{
"name": "textInput",
"description": "Single-line text.",
"arg": "default_value",
"notes": {
"value": {
"visible": true,
"required": false,
"type": "string",
"default": "default_value"
},
"details": {
"label": "Text Input"
}
}
}
```
If a `notes.details.label` field is not present,
then the DE will use the `name` as the parameter's form field label.
The `description` will be used as the help text displayed below the form field in the DE.
Note that the `visible`, `required`, and `type` fields could be omitted from the `notes` in this example,
since those are the defaults the DE will use for those parameter fields.
This text field will render like this in the DE's app launch form:
{ width="464" }
Text Form Field
If the parameter has an `inputMode` field with a `REQUIRED` value,
then the DE will treat that parameter as required and ignore the
`notes.value.required` field.
If the parameter has an `inputMode` field with a `FIXED` value,
then the DE will treat that parameter as hidden and ignore the
`notes.value.visible` field.
The `notes` object can contain any other custom fields,
which will be ignored by the DE,
except for certain fields in specific parameter types,
detailed in the following sections.
### Parameter Types
#### Inputs
If an input value is also required as a parameter on the command line,
then the `name` of the input in the `fileInputs` parameter field
should match the `name` of the corresponding `appArgs` parameter (or the `key` of the corresponding `envVariables` parameter).
The `requiredInput` field from the example above will render like this in the DE's app launch form:
{ width="464" }
File/Folder Input Field
If the app's `runtime` value is `DOCKER`,
then the DE will automatically prepend `/TapisInput/`
to the file or folder name submitted by the user,
since that is the directory mounted inside the container by Tapis for inputs.
Otherwise only the base name of the file or folder will be submitted
for the command line argument.
#### Outputs
The `outputs` field is not used by Tapis, but if a user wishes to create a
pipeline workflow with multiple Tapis and DE apps,
then the DE requires an app output to be defined so that it can be used as an
app's input in a subsequent step in the workflow.
If the output value is also required as a parameter on the command line,
then the `name` of the output in the `notes` field should match the `name` of the
corresponding `appArgs` parameter (or the `key` of the corresponding `envVariables` parameter).
The `requiredOutput` field from the example above will render like this in the DE's app launch form:
{ width="464" }
Output Field
If the app's `runtime` value is `DOCKER`,
then the DE will automatically prepend `/TapisOutput/`
to the parameter value submitted by the user,
since that is the directory mounted inside the container by Tapis for outputs.
Otherwise the DE will automatically prepend `output/`
to the parameter value submitted by the user,
since that is the directory created and used by Tapis for app outputs.
#### Flags or Booleans
If the parameter has a `notes.value.type` field value of `bool`, `boolean`, or `flag`,
then the DE will display that parameter as a checkbox.
Parameters with a `notes.semantics.ontology` list where the first value is
`xs:boolean` will also display as a checkbox.
Note that the value of a parameter's `arg` is used in job submissions
when the user checks a flag parameter type in the submission form.
The `flag` fields from the example above will render like this in the DE's app launch form:
{ width="464" }
Flag Input Fields
#### Enumeration Lists
If the parameter has a `notes.value.type` field value of `enumeration`,
then the DE will display that parameter as a selection list.
The parameter should also have a `notes.value.enum_values` field
with an array of object values.
Each of these enum objects should have a key
that is the full command line argument that should be submitted as the parameter's `arg` value in the job submission,
and a value that will be used as the display label by the DE in the form's selection list.
The `listInput` field from the example above will render like this in the DE's app launch form:
{ width="464" }
List Selection Field
{ width="464" }
List Selection Field Expanded
#### Numbers
If the parameter has a `notes.value.type` field value of `number`,
then the DE will display that parameter as a number field.
The DE will use the first `xs:*` value in the `notes.semantics.ontology` field to
determine if the user's input should be restricted to integers or decimal values.
See https://github.com/cyverse-de/mescal/blob/main/src/mescal/tapis_de_v3/params.clj
for a list of supported number XSD types (otherwise defaulting to decimal values).
The `integerInput` and `doubleInput` fields from the example above will render like this in the DE's app launch form:
{ width="464" }
Number Input Fields
### Make App Public
Once the app is tested and ready for sharing with all DE users,
you can make the app public with the
[/v3/apps/share_public endpoint](https://tapis-project.github.io/live-docs/?service=Apps#tag/Sharing/operation/shareAppPublic){target=_blank}:
curl-tapis -s -X POST "https://cyverse.tapis.io/v3/apps/share_public/CyVerse-QA-Test-App"
!!! tip "UNM CARC users"
Discovery Environment HPC apps run on TACC systems through Tapis. UNM
researchers also have CARC's own clusters: CyVerse is one of
[CARC's partner platforms](https://carc.unm.edu/docs/about/partners/){target=_blank}, and
[Transferring data](https://carc.unm.edu/docs/getting-started/transferring-data/){target=_blank} explains how to move inputs
and results between CARC storage and other systems.
!!! info "Under the hood"
Migrating older Tapis v2 (Agave) apps to Tapis v3 for the Discovery Environment is covered in [Tapis v2 to v3 migration](https://docs.cyverse.org/api/tapis-v2-v3-migration/){target=_blank} in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/create_hpc_apps.md){target=_blank} (last source update 2026-03-16), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/discovery-environment/vice/overview/
---
title: "Interactive analysis in the Discovery Environment"
description: "VICE interactive apps in the Discovery Environment: featured apps, requesting VICE access, launching, Data Store access, Git, and instant and quick launches."
type: Guide
tags:
- VICE
- Interactive Apps
- Discovery Environment
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/vice/about.md"
title: "CyVerse Learning Materials: docs/de/vice/about.md"
author: "team:cyverse"
last_modified: "2025-03-26T11:21:50-07:00"
---
# Interactive analysis in the Discovery Environment
**VICE** stands for Visual Interactive Computing Environment and is a part of CyVerse's [Discovery Environment (DE)](https://de.cyverse.org){target=_blank}.
CyVerse maintains featured apps from the [Rocker-Project](https://rocker-project.org/images/){target=_blank}, [Project Jupyter](https://jupyter-docker-stacks.readthedocs.io/en/latest/index.html){target=_blank}, and [Visual Studio Code](https://code.visualstudio.com/docs/remote/create-dev-container){target=_blank}.
Our Docker images are built from community-maintained image registries (i.e., [DockerHub](https://hub.docker.com/){target=_blank}, [GitHub Container Registry](https://docs.github.com/en/packages/working-with-a-github-packages-registry/working-with-the-container-registry){target=_blank}, [NVIDIA GPU Container Registry](https://catalog.ngc.nvidia.com/){target=_blank}, [BioContainers](https://biocontainers.pro/){target=_blank}, with a few additonal packages for use in CyVerse DE.
There are a few common categories of featured interactive applications:
**Linux Shell**
- To add this badge to your own `README.MD`:
```html
```
**Integrated Development Environments (IDE)**
[:simple-rstudioide: Rocker RStudio](https://rocker-project.org/){target=_blank}
- To add this badge to your own `README.MD`:
```html
```
[:simple-jupyter: Project Jupyter Lab](https://quay.io/organization/jupyter){target=_blank}
- To add this badge to your own `README.MD`:
```html
```
[:material-microsoft-visual-studio-code: Visual Studio Code Server](https://code.visualstudio.com/docs/remote/vscode-server){target=_blank}
- To add this badge to your own `README.MD`:
```html
```
**Virtual Desktop Environments**
[KASM Desktops](https://hub.docker.com/r/kasmweb/desktop){target=_blank}
- To add this badge to your own `README.MD`:
```html
```
**Web Server Applications**
- StreamlitApps, ShinyApps, WebGL, HTML5, etc.
CyVerse hosts the container recipes (Dockerfiles) of its featured apps on GitHub: . These images are maintained by CyVerse staff.
??? tip "Getting VICE Access"
Since VICE (the Visual Interactive Computing Environment) is a target for cryptocurrency miners, we require an additional verification.
To get access to VICE, your CyVerse account must be associated with a valid email address from an organization, an educational institution, or a government; e.g., an email ending with `.org`, `.edu`, or `.gov`. We do not grant VICE access to commercial email addresses, e.g., `@gmail.com`, `@yahoo.com`, `@msn.com`, etc.
To request VICE access, visit the [User Portal](https://user.cyverse.org/services){target=_blank} and **Services**; look for [![][vice]{width=30}](https://user.cyverse.org/services){target=_blank} [DE VICE](https://user.cyverse.org/services){target=_blank} and select the **REQUEST ACCESS** link.
You will be asked for a description of your intended VICE use: please give non-technical scientific details, and if you can, link an external resource (like a workshop or lab website) and funding agency.
-----------------------------------------------------------------------
## Launching Applications
When you open the Apps Tab in the DE, there are several featured Applications immediately visible, these are all Interactive Apps.
These apps launch with their default number of cores, amount of RAM, and timeout, and without input data.
??? tip "Instant Launching Apps"
[![][vice]{width=60}](https://de.cyverse.org/instantlaunches){target=_blank}
Pre-configured apps with base settings can be launched with one click from the Instant Launch menu.
These apps have 4-cores and 16 GB RAM pre-set limit.
1. Log into the [![][de]{width=25}](https://de.cyverse.org){target=_blank} [Discovery Environment](https://de.cyverse.org){target=_blank}.
2. Click on the [![][apps]{width=25}](https://de.cyverse.org/apps){target=_blank} [Apps](https://de.cyverse.org/apps){target=_blank}.
3. Select one of the Featured Apps.
4. Adjust the following:
Under "Analysis Info", for **Output Folder** click **Browse** and navigate to and select your new `/analysis/` folder and re-name it, or leave it as the default name, then click **Next**;
In **Advanced Settings** you can modify the resources used by the container (e.g., more cores, more RAM, or disk space for larger analyses) or leave the default settings, then click **Next**;
Click **Launch Analysis** to launch your application.
Your analysis should be available within one minute. If it takes more than five minutes, terminate the analysis and try again.
5. In the navigation bar, click on the [![][analyses]{width=30}](https://de.cyverse.org/analyses){target=_blank} [Analyses](https://de.cyverse.org/de/analyses){target=_blank} view.
Your application will first appear as "Submitted" (this takes a minute or two; maybe more depending on both the size of the container and any additional imported data).
When the Status of the launch is "Running", click on the running App under "Analyses"; a new tab will open in your browser.
??? tip "Be patient, but not too patient with custom apps or large images"
Featured images are cached on the resource nodes and should start in under 30 seconds.
However, newly updated images must be pulled from our container registry or from public registries.
Data Science and Machine Learning images can be two to twenty gigabytes (Gb), and may take five to ten minutes to download and inflate the first time they are run.
Even when the application has entered 'Running' status, it may take additional time for input data to be transferred onto the resource with the new container.
-----------------------------------------------------------------------
## Accessing data from VICE apps
VICE apps have web access, so you can import data using methods like `curl` or querying external databases.
You can also access your Data Store `/iplant/home/` and `/iplant/home/shared/` directories directly from VICE apps via the CSI Driver.
The Data Store is a CSI driver mount at `/data-store`; `~/data-store` is a symbolic link to it, so either path works. Your Data Store home folder is at `/data-store/iplant/home//` (also `~/data-store/iplant/home//`). Public Projects and private Projects shared with your username are under `/data-store/iplant/home/shared/`. See [Data Store in interactive apps](https://unm-carc.github.io/cyverse/discovery-environment/vice/csi/) for the full mount layout.
You will be able to read, write, and delete data from the Data Store from a VICE app, just as you would on a local filesystem.
**Important: data under `/data-store` (including `~/data-store`) is read over the CSI driver and has much slower I/O performance than the container's local disk, such as your home directory `~/`. Copy large or many-file working sets to local disk while you work.**
!!! info "CSI Driver in VICE"
VICE Apps takes advantage of the [Kubernetes Container Storage Interface (CSI) Driver](https://kubernetes-csi.github.io/docs/introduction.html){target=_blank}.
This feature brings your Data Store into the container with you every time you launch a new analysis.
Remember: these mounted data are being viewed over the network and when they are moved or modified their performance is much slower than the data physically located on the hard disk drives or SSD of the host where your Analysis container is running.
In general the CSI driver can handle Notebooks and small data files without any noticeable differences.
You will begin to see degradation in performance for operations requiring access to many files or very large files.
For these types of processes, we recommend making a copy of your data in your current working directory, and moving them back to the Data Store when you're finished.
Data can be moved over the CSI driver using normal UNIX commands like `cp` and or `mv` but be aware that any modifications you make will be recorded on the Data Store.
Files that you have `read-only` access to will not be modified on the Data Store.
??? tip "Working with many files"
Read/write speeds for single files are quite fast, but can slow down when accessing many files.
If you are working with many small files, it may be faster to keep your data in the Data Store in a compressed format (such as .tgz or .zip), then use `cp` to copy the data from `~/work/home/username/` to `~/work`. Working with many small files within the VICE app's container will be faster than accessing them directly from the Data Store.
Speed benchmark for moving a folder with many CSV files:
moving storms_by_year/ folder: 23.5s
moving storms_by_year.zip: 0.07s
unzipping: 0.063s
-----------------------------------------------------------------------
## Using Git and GitHub from VICE
Git and Github are essential tools for software version control. The following documentation show how to use Git and Github with the Cyverse Discovery Environment and specifically with Visual Interactive Computing Environement (VICE) apps (e.g, Jupyter Lab). The core VICE apps (RStudio, JupyterLab, and Cloud Shell) have the [`GitHub command line interface`][gh] installed.
### The Problem
Your personal Data Store directory (i.e., `~/data-store/iplant/home/`) would seem like the most logical place to clone git repositories, work in them, then push changes back up to Github. Unforntunately, git repositories are not compatable with Integrated Rule Oriented Data System (IRODS) which is the underlying technology of the Cyverese Datastore. So to use git and github with Cyverse, we need a different solution.
### The Solution
Instead of doing git commands from your personal datastore, we can do git from the home directories of JupyterLab or Cloudshell containers. When you first launch a JupyterLab terminal, you will probably be in your home directory `~/`, which is on the container's local disk; git commands work perfectly fine there. (`~/data-store` is a link to the `/data-store` CSI mount, so it is not local disk.) To do git commands with Github, we need to have `.gitconfig` and `.ssh` files stored in our container. A problem with this is that Cyverse containers are ephemeral and disappear when the App is shutdown. That means every time a CyVerse Jupyter App is launched, users need to create git credentials and an ssh key for a GitHub handshake. Quite annoying!
We can work around this issue by creating `.config` and `.ssh` files one time in the container and then store them in your personal Datastore directory. Each time you start a new JupyterLab container we need to copy the `.config` and `.ssh` files from your personal Datastore directory back to the container.
#### Step 1: Set up the authentication handshake with Github
- Launch a VICE App with Jupyter (e.g., [JupyterLab Data Science](https://de.cyverse.org/apps/de/c2227314-1995-11ed-986c-008cfa5ae621){target=_blank})
- Once the App is running, open the App's terminal.
- `ssh`:
1. Create your ssh key with `ssh-keygen`. Use whichever encryption you prefer.
2. Copy your `ssh` folder to your own private CyVerse Data Store folder (`cp -r ~/.ssh/ ~/data-store/iplant/home//`)
3. Copy your publish `ssh` key to GitHub (print key with `cat ~/.ssh/id_rsa.pub`, copy the key, go to https://github.com/settings/keys and add your key)
- `git` credentials:
1. Create your `git` credentials with `git config --global user.name ""` and `git config --global user.email ""`
2. Copy your `git`credentials to your own private CyVerse Data Store folder (`cp ~/.gitconfig ~/data-store/iplant/home//`)
You can now work on your cloned GitHub repositories and push changes.
#### Step 2: Reproducing the handshake
Likely, once done with work on the launched App, you will terminate it and therefore lose the `ssh` and `git` credentials.
Since we have copied the necessary files to our own private Data Store folders, we can reproduce this handshake by just copying these files to any newely launched App.
1. Copy the `ssh` files with `cp -r ~/data-store/iplant/home//.ssh ~/`
2. Copy the `git` credentials with `cp ~/data-store/iplant/home//.gitconfig ~/`
You should now be able to work on your cloned GitHub repositories and push changes without having to recreate the `ssh` key or your `git` credentials.
!!! tip
Remember that as you are working in a git repository in a Cyverse container, the container is ephemeral and will disappear when the App is shutdown. You will also lose any changes you made in the repository. PLEASE REMEMBER TO PUSH CHANGES TO GITHUB OFTEN! Cyverse containers will self-destruct after the alloted time.
!!! tip "Working with Git repositories"
For the time being, we recommend cloning repositories into `~/` (the container's local disk) rather than anywhere under `/data-store` or `~/data-store`, because the large number of files in a Git repository can make transfers to the Data Store slow. We are working on optimizing the `git clone` process to address this issue.
-----------------------------------------------------------------------
## Instant Launches
From the Home tab in the DE, there are several apps that have an **Instant Launch** feature which allows you to start the app with a single click.
These apps launch with their default number of cores, amount of RAM, and timeout, and without input data. You can always import data using HTTPS protocols or iCommands after launch.
-----------------------------------------------------------------------
## Quick Launches
Quick launch buttons are directed URLs that allow you to share an app with pre-set configurations. After selecting an app, you will be taken to the app launcher where you can select input data sets, and then set size and time parameters. Any public app can have a quick launch URL generated for it.
To use one of the Featured Launches listed below in the table, copy the badge (button link) to add to wherever you collaborate (a webpage, project notes, documentation, etc.).
To create your own Saved Launch, start by launching the app you want to use. This will be a good time to Favorite the app. In the "Review & Launch" panel, click the "Create Saved Launch" button. You will be asked to name your Saved Launch and check the box when prompted if you would like it to be public. Remember which app you saved, you will find the link under Details of the app you saved.
[gh]: https://cli.github.com/
[gh-manual]: https://cli.github.com/manual/index
[de]: https://unm-carc.github.io/cyverse/assets/de/logos/deIcon.svg
[home]: https://unm-carc.github.io/cyverse/assets/de/menu_items/homeIcon.svg
[data]: https://unm-carc.github.io/cyverse/assets/de/menu_items/dataIcon.svg
[apps]: https://unm-carc.github.io/cyverse/assets/de/menu_items/appsIcon.svg
[analyses]: https://unm-carc.github.io/cyverse/assets/de/menu_items/analysisIcon.svg
[vice]: https://unm-carc.github.io/cyverse/assets/de/logos/deviceIcon.png
!!! info "Under the hood"
How VICE analyses are deployed on Kubernetes and given Data Store access through the iRODS CSI driver is described on the [VICE](https://docs.cyverse.org/deployment/06-applications/vice/){target=_blank} deployment page in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/vice/about.md){target=_blank} (last source update 2025-03-26), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/discovery-environment/vice/quick-rstudio/
---
title: "RStudio in 6 steps"
description: "Quick start: launch RStudio as a VICE interactive app, create a project, and terminate the app when you finish."
type: Tutorial
tags:
- RStudio
- VICE
- Quick Start
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/vice/quick-rstudio.md"
title: "CyVerse Learning Materials: docs/de/vice/quick-rstudio.md"
author: "team:cyverse"
last_modified: "2025-03-16T09:25:29-07:00"
---
# RStudio in 6 steps
## 1. Log into Discovery Environment
Log into
If you have not yet created an account, go to the [User Portal](https://user.cyverse.org){target=_blank} and sign up.
## 2. Launch the App
[![rstudio_1]][rstudio_1]
[rstudio_1]: https://unm-carc.github.io/cyverse/assets/de/rstudio_1.png
Click on the **Apps** grid icon
[RStudio Verse](https://de.cyverse.org/apps/de/3548f43a-bed1-11e9-af16-008cfa5ae621/launch){target=_blank} is in "Featured Apps".
[Instant Launches](https://de.cyverse.org/instantlaunches){target=_blank} start Apps immediately when clicked.
The conventional launch menu allows you to modify the App parameters. You can add input data, increase the amount of RAM or CPU cores, and change the analysis directory.
[](https://de.cyverse.org/apps/de/3548f43a-bed1-11e9-af16-008cfa5ae621/launch)
## 3. Open the Analysis
After you have started a VICE app, a new tab will automatically open in your browser and take you to the loading screen.
[![rstudio_3]][rstudio_3]
[rstudio_3]: https://unm-carc.github.io/cyverse/assets/de/rstudio_3.png
Once the app is ready, it will transition to the user interface.
[![rstudio_4]][rstudio_4]
[rstudio_4]: https://unm-carc.github.io/cyverse/assets/de/rstudio_4.png
**RStudio Interface:**
RStudio is a free, open source IDE (integrated development environment) for R.
Its interface is organized so that the user can clearly view graphs, data tables, R code and ouput at the same time.
It also offers an Import-Wizard-like feature that allows users to import CSV, Excel, SAS (*.sas7bdat), SPSS (*.sav), and Stata (\*.dta) files into R without having to write the code to do so.
More information about RStudio can be found [here](https://www.rstudio.com/products/rstudio/){target=_blank}.
!!! note "Long wait times?"
Normal wait time for a featured VICE app to launch is less than 2 minutes. If you're experiencing a significantly longer wait, consider terminating the Analysis and starting a new one.
## 4. Create an RStudio Project
You can create RStudio projects using local data, or from Git.
[![rstudio_5]][rstudio_5]
[rstudio_5]: https://unm-carc.github.io/cyverse/assets/de/rstudio_5.png
This example uses [Leaflet Maps](https://github.com/rstudio/leaflet){target=_blank} in RStudio.
[![rstudio_6]][rstudio_6]
[rstudio_6]: https://unm-carc.github.io/cyverse/assets/de/rstudio_6.png
You can then run R commands and install packages.
[![rstudio_7]][rstudio_7]
[rstudio_7]: https://unm-carc.github.io/cyverse/assets/de/rstudio_7.png
[![rstudio_8]][rstudio_8]
[rstudio_8]: https://unm-carc.github.io/cyverse/assets/de/rstudio_8.png
## 5. Terminate your app
The Discovery Environment is a shared system. In fairness to the community, users should "Terminate" any apps that
are no longer actively running analyses.
In the Analyses window, select the app (by clicking the checkbox next to it), then select "More Actions", then "Terminate" to shut down the app.
[![rstudio_9]][rstudio_9]
[rstudio_9]: https://unm-carc.github.io/cyverse/assets/de/rstudio_9.png
Any new data in the `/home/rstudio/data-store/data/output` directory will begin copying back to your folder at this time.
Any input data which you added when the app started using the conventional launch feature will *not* be copied.
!!! warning "Automatic Termination and Extension"
VICE apps run for a pre-determined amount of time, typically between 4 and 48 hours.
If you have opted for email notifications from the DE, then you'll get a notification 1 day before and another 1 hour before the app will terminate.
To extend the pre-set run time, go to your analysis and click the hour glass icon which automatically extends the app run time.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/vice/quick-rstudio.md){target=_blank} (last source update 2025-03-16), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/discovery-environment/vice/quick-jupyter/
---
title: "JupyterLab in 6 steps"
description: "Quick start: launch JupyterLab as a VICE interactive app, work with Data Store files, and terminate the app when you finish."
type: Tutorial
tags:
- JupyterLab
- VICE
- Quick Start
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/vice/quick-jupyter.md"
title: "CyVerse Learning Materials: docs/de/vice/quick-jupyter.md"
author: "team:cyverse"
last_modified: "2025-03-16T09:25:29-07:00"
---
# JupyterLab in 6 steps
## 1. Log into Discovery Environment
Log into .
If you have not yet created an account, go to the [User Portal](https://user.cyverse.org){target=_blank} to sign up.
## 2. Launch the App
[![jupyter_1]][jupyter_1]
[jupyter_1]: https://unm-carc.github.io/cyverse/assets/de/jupyter_1.png
Click on the **Apps** grid icon.
[Jupyter Lab Datascience](https://de.cyverse.org/apps/de/cc77b788-bc45-11eb-9934-008cfa5ae621/launch){target=_blank} is in "Featured Apps".
[Instant Launches](https://de.cyverse.org/instantlaunches){target=_blank} start Apps immediately when clicked.
The conventional launch menu allows you to modify the App parameters. You can add input data, increase the amount of RAM or CPU cores, and change the analysis directory.
[![jupyter_2]][jupyter_2]
[jupyter_2]: https://unm-carc.github.io/cyverse/assets/de/jupyter_2.png
## 3. Open the Analysis
After you have started a VICE app, a new tab will automatically open in your browser and take you to the loading screen.
[![jupyter_3]][jupyter_3]
[jupyter_3]: https://unm-carc.github.io/cyverse/assets/de/jupyter_3.png
[![jupyter_4]][jupyter_4]
[jupyter_4]: https://unm-carc.github.io/cyverse/assets/de/jupyter_4.png
Once the app is ready, it will transition to the user interface.
[![jupyter_5]][jupyter_5]
[jupyter_5]: https://unm-carc.github.io/cyverse/assets/de/jupyter_5.png
**The Jupyter Lab Interface:**
While Jupyter Lab has many features found in traditional integrated development environments (IDEs), it remains focused on interactive, exploratory computing.
The Jupyter Lab interface consists of a main work area containing tabs of documents and activities, a collapsible left sidebar, and a menu bar.
The left sidebar contains a file browser, the list of running kernels and terminals, the command palette, the notebook cell tools inspector, and the tabs list.
More information about the Jupyter Lab can be found [here](https://jupyterlab.readthedocs.io/en/stable/user/interface.html){target=_blank}.
!!! note "Long wait times?"
Normal wait time for a featured VICE app to launch is less than 2 minutes. If you're experiencing a significantly longer wait, consider terminating the Analysis and starting a new one.
## 4. Create a new `conda` environment
From Jupyter's Launch menu, select the black Terminal console icon.
This will take you to a command line shell.
Change directory, or download a sample `environment.yml` file:
```bash
$ cd /home/shared/cyverse_training/platform_guides/discovery_environment/jupyterlab/
$ conda env create -f environment.yml
```
or
```bash
$ wget https://data.cyverse.org/dav-anon/iplant/commons/community_released/cyverse_training/platform_guides/discovery_environment/jupyterlab/environment.yml
$ conda env create -f environment.yml
```
and then:
```bash
$ conda activate python39
```
## 5. Create Jupyter notebook
Jupyter notebooks (`.ipynb`) combine code with narrative text (Markdown), equations (LaTeX), images and interactive visualizations.
To create a notebook, click the `+` button which opens the new Launcher tab.
The JupyterLab Datascience containers have three pre-installed kernels: Python3, Julia, and R.
[Official Jupyter Notebooks](https://jupyterlab.readthedocs.io/en/stable/user/notebook.html){target=_blank}
To open the classic Notebook view from JupyterLab, select "Launch Classic Notebook" from the Help Menu.
## 6. Terminate your app
The Discovery Environment is a shared system. In fairness to the community, users should "Terminate" any apps that
are no longer actively running analyses.
In the Analyses window, select the app (by clicking the checkbox next to it), select "More Actions", then "Terminate" to shut down the app.
[![jupyter_7]][jupyter_7]
[jupyter_7]: https://unm-carc.github.io/cyverse/assets/de/jupyter_7.png
Any new data in the `/home/jovyan/work/data/outputs` directory will begin copying back to your folder at this time.
Any input data which you added when the app started using the conventional launch feature will *not* be copied.
!!! warning "Automatic Termination and Extension"
VICE apps run for a pre-determined amount of time, typically between 4 and 48 hours.
If you have opted for email notifications from the DE, then you'll get a notification 1 day before and another 1 hour before the app will terminate.
To extend the pre-set run time, go to your analysis and click the hour glass icon which automatically extends the app run time.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/vice/quick-jupyter.md){target=_blank} (last source update 2025-03-16), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/discovery-environment/vice/quick-cloudshell/
---
title: "Cloud Shell in 6 steps"
description: "Quick start: launch the Cloud Shell terminal app in the Discovery Environment and use it to work with the Data Store."
type: Tutorial
tags:
- Cloud Shell
- VICE
- Quick Start
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/vice/quick-cloudshell.md"
title: "CyVerse Learning Materials: docs/de/vice/quick-cloudshell.md"
author: "team:cyverse"
last_modified: "2025-03-16T09:25:29-07:00"
---
# Cloud Shell in 6 steps
## 1. Log into the Discovery Environment
Log into
If you have not yet created an account, go to the [User Portal](https://user.cyverse.org){target=_blank} and sign up.
## 2. Launch the App
The Cloud Shell icon is found in the left-side navigation bar in the Discovery Environment.
[![cloud_shell_1]][cloud_shell_1]
[cloud_shell_1]: https://unm-carc.github.io/cyverse/assets/de/cloud_shell_1.png
Click on the **Apps** grid icon
[Cloud Shell](https://de.cyverse.org/apps/de/5f2f1824-57b3-11ec-8180-008cfa5ae621/launch){target=_blank} is listed at the top of "Featured Apps".
[Instant Launches](https://de.cyverse.org/instantlaunches){target=_blank} start Apps immediately when clicked.
The conventional launch menu allows you to modify the App parameters. You can add input data, increase the amount of RAM or CPU cores, and change the analysis directory.
[![cloud_shell_2]][cloud_shell_2]
[cloud_shell_2]: https://unm-carc.github.io/cyverse/assets/de/cloud_shell_2.png
## 3. Open the Analysis
After you have started a VICE app, a new tab will automatically open in your browser and show the app's loading screen.
[![cloud_shell_3]][cloud_shell_3]
[cloud_shell_3]: https://unm-carc.github.io/cyverse/assets/de/cloud_shell_3.png
Once the app is ready, it will transition to the user interface (in this example, a Linux terminal).
[![cloud_shell_4]][cloud_shell_4]
[cloud_shell_4]: https://unm-carc.github.io/cyverse/assets/de/cloud_shell_4.png
You should see a "message of the day", information about the machine you're using, and your CyVerse username for when you initiate
an iCommands connection.
!!! note "Long wait times?"
Normal wait time for a featured VICE app to launch is less than 2 minutes. If you're experiencing a significantly longer wait, consider terminating the Analysis and starting a new one.
## 4. Activate a `conda` environment
The Cloud Shell comes with multiple languages and package managers pre-installed. These include `go`, `python`, and `rust`.
Package managers include `conda` and `cargo`, in addition to linux `apt` and `apt-get`.
Your identity inside the container is `user` and you have limited `sudo` privileges to install new packages into the container.
These changes are not saved after the analyses ends, or when you start a new Cloud Shell Analysis later.
!!! tip "Behind the CLI engine"
The Cloud Shell is running a terminal multiplexer called [tmux](https://en.wikipedia.org/wiki/Tmux){target=_blank} which keeps your session active even after you've closed your browser tab. Here is a [cheat sheet](https://tmuxcheatsheet.com/){target=_blank} that can help you with tmux!
[`tmux` key bindings](http://manpages.ubuntu.com/manpages/bionic/man1/tmux.1.html) are active.
One side effect of using `tmux` is that you cannot scroll up in the Terminal to see previous outputs. This can make it difficult to view long outputs. If your output doesn't fit on the screen and you still want to see the whole thing, you can pipe the results of the command into `pager`.
For instance, if you run `head big_file.csv -n 100` to view the first 100 lines of a CSV, it probably won't all fit on screen. If you run `head big_file.csv -n 100 | pager`, you will be able to move through the entire output. In the `pager`, you use the J key to scroll down, the K key to scroll up, and the Q key to exit.
To activate conda:
```
conda init
```
and then:
```
conda activate base
```
If you receive a message about refreshing your screen, you can `exit` the Cloud Shell by typing "exit" and then clicking the refresh button on your browser tab.
## 5. Using `icommands`
To connect to the CyVerse Data Store, you can initiate an iRODS iCommands `iinit`.
You should now be connected to your `/iplant/home/username` home directory.
### ils
```
ils /iplant/home/username/
```
To view the 'shared' directory, type:
```
ils /iplant/home/shared
```
### iget
Download data into your Cloud Shell with [iCommands](https://docs.irods.org/master/icommands/user/){target=_blank} by running `iget`:
``` iget -KPbvrf /iplant/home/shared/cyverse_training/ ```
### iput
After you finish your analyses, you can save the outputs to Data Store.
Use `iput` to copy your new files back to your user space, or if you've left your new work in the `/home/user/work/data/outputs` folder, it will be copied back to your `/iplant/home/username/Analyses/` directory.
To find the outputs you generated (if any), use the same steps as before, but this time select the 'Go To Output Folder'.
## 6. Terminate your app
The Discovery Environment is a shared system. In fairness to the community, users should "Terminate" any apps that
are no longer actively running analyses.
In the Analyses window, select the app (by clicking the checkbox next to it), select "More Actions", and then "Terminate" to shut down the app.
Any new data in the `/home/user/work/data/output` directory will begin copying back to your folder at this time.
Any input data which you added when the app started using the conventional launch feature will *not* be copied.
!!! Warning "Automatic Termination and Extension"
VICE apps run for a pre-determined amount of time, typically between 4 and 48 hours.
If you have opted for email notifications from the DE, then you'll get a notification 1 day before and another 1 hour before the app will terminate.
To extend the pre-set run time, go to your analysis and click the hour glass icon which automatically extends the app run time.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/vice/quick-cloudshell.md){target=_blank} (last source update 2025-03-16), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/discovery-environment/vice/quick-ollama/
---
title: "Running your LLM for inference with Ollama"
description: "Quick start: run open-source large language models with Ollama inside a GPU-enabled JupyterLab PyTorch app in the Discovery Environment."
type: Tutorial
tags:
- Ollama
- LLM
- GPU
- VICE
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/vice/quick-ollama.md"
title: "CyVerse Learning Materials: docs/de/vice/quick-ollama.md"
author: "team:cyverse"
last_modified: "2025-03-16T09:25:29-07:00"
---
# Running your LLM for inference with Ollama
## 1. Log into Discovery Environment
Log into .
If you have not yet created an account, go to the [User Portal](https://user.cyverse.org){target=_blank} to sign up.
## 2. Launch the JupyterLab Pytorch GPU App
[![ollama_1]][ollama_1]
[ollama_1]: https://unm-carc.github.io/cyverse/assets/de/ollama_1.png
Click on the **Apps** grid icon.
[Jupyter Lab Pytorch GPU](https://de.cyverse.org/apps/de/e19a5772-94e6-11ec-b1f0-008cfa5ae621/launch){target=_blank} is in "Featured Apps".
[Instant Launches](https://de.cyverse.org/instantlaunches){target=_blank} start Apps immediately when clicked.
The conventional launch menu allows you to modify the App parameters. You can add input data, increase the amount of RAM or CPU cores, and change the analysis directory.
[![ollama_2]][ollama_2]
[ollama_2]: https://unm-carc.github.io/cyverse/assets/de/ollama_2.png
## 3. Open the Analysis
After you have started a VICE app, a new tab will automatically open in your browser and take you to the loading screen.
[![ollama_3]][ollama_3]
[ollama_3]: https://unm-carc.github.io/cyverse/assets/de/ollama_3.png
[![ollama_4]][ollama_4]
[ollama_4]: https://unm-carc.github.io/cyverse/assets/de/ollama_4.png
Once the app is ready, it will transition to the user interface.
[![ollama_5]][ollama_5]
[ollama_5]: https://unm-carc.github.io/cyverse/assets/de/ollama_5.png
We will focus on running an llm for inference using _Ollama as a server_. The Jupyter Lab with Pytorch GPU, comes with Ollama preinstalled, which is necessary to run LLMs.
The Jupyter Lab interface consists of a main work area containing tabs of documents and activities, a collapsible left sidebar, and a menu bar.
The left sidebar contains a file browser, the list of running kernels and terminals, the command palette, the notebook cell tools inspector, and the tabs list.
!!! note "Long wait times?"
Normal wait time for a featured VICE app to launch is less than 2 minutes. If you're experiencing a significantly longer wait, consider terminating the Analysis and starting a new one.
## 4. Download and run an LLM using Ollama
From Jupyter's Launch menu, select the black Terminal console icon.
This will take you to a command line shell.
Run the following command to start Ollama in _server mode_:
```bash
$ ollama serve
```
Without closing the shell tab, Ollama will be running in server mode, ready to be called from any jupyter notebook or python script.
Open a new tab with a console and then, pull the model you're interested in. For example `llama 3.1`:
```bash
$ ollama pull llama3.1
```
It is necessary to fetch ahead of time the model you're interested in using. You can find the list of supported LLMs at (https://ollama.com/search)[Ollama's Hub].
## 5. Create Jupyter notebook to interface with the LLM.
[Ollama LLM usage example](https://github.com/ua-datalab/Generative-AI/blob/main/Notebooks/Running%20LLM%20locally%20-%20Ollama.ipynb){target=_blank}
After starting Ollama in server mode and downloading the relevant models, you can run your code that uses the aformentioned models.
To create a notebook, click the `+` button which opens the new Launcher tab.
To open the classic Notebook view from JupyterLab, select "Launch Classic Notebook" from the Help Menu.
Create a new jupyter notebook to where can use any llm client library to interface with the LLM. DataLab has a reference notebook with a usage example using ollama's client libary.
Alternatively, you can use [langchain](https://python.langchain.com/docs/introduction/){target=_blank}, [llama-index](https://docs.llamaindex.ai/en/stable/){target=_blank}, the [OpenAI API client library](https://github.com/openai/openai-python){target=_blank} or any other client library that speaks OpenAI's REST API.
## 6. Terminate your app
The Discovery Environment is a shared system. In fairness to the community, users should "Terminate" any apps that
are no longer actively running analyses.
In the Analyses window, select the app (by clicking the checkbox next to it), select "More Actions", then "Terminate" to shut down the app.
[![ollama_7]][ollama_7]
[ollama_7]: https://unm-carc.github.io/cyverse/assets/de/ollama_7.png
Any new data in the `/home/jovyan/work/data/outputs` directory will begin copying back to your folder at this time.
Any input data which you added when the app started using the conventional launch feature will *not* be copied.
!!! warning "Automatic Termination and Extension"
VICE apps run for a pre-determined amount of time, typically between 4 and 48 hours.
If you have opted for email notifications from the DE, then you'll get a notification 1 day before and another 1 hour before the app will terminate.
To extend the pre-set run time, go to your analysis and click the hour glass icon which automatically extends the app run time.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/vice/quick-ollama.md){target=_blank} (last source update 2025-03-16), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/discovery-environment/vice/csi/
---
title: "Data Store in interactive apps"
description: "How the Kubernetes CSI driver mounts the Data Store at /data-store inside VICE apps, its latency limits, and when to use it."
type: Guide
tags:
- VICE
- Data Store
- Kubernetes
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/vice/csi.md"
title: "CyVerse Learning Materials: docs/de/vice/csi.md"
author: "team:cyverse"
last_modified: "2025-03-14T14:47:33-07:00"
---
# Data Store in interactive apps
## Overview
CyVerse Discovery Environment uses Kubernetes to orchestrate interactive applications including Jupyter, RStudio, Remote Desktops, etc.
The Kubernetes Container Storage Interface (CSI) is a standardized interface that allows Kubernetes to interact with a wide variety of storage systems. By utilizing CSI, Kubernetes can dynamically provision, attach, mount, and manage storage volumes across different storage providers without being tightly coupled to any specific storage technology. This flexibility is crucial for ensuring that your containerized applications have reliable and consistent access to the storage they require, regardless of the underlying infrastructure.
## Purpose in the Discovery Environment VICE Apps
In our workbench environment, the CSI driver plays a key role in managing data in the Data Store.
We have integrated the CSI driver to facilitate seamless access and management of user and community data stored in the Data Store within containerized applications.
The driver creates volume mounts at specific paths in our containers:
`/data-store`: A root of the volume mounts for accessing community shared and personal data in the Data Store. This directory contains following four sub-directories.
`/data-store/input`: A directory providing access to the input data specified when launching the app. If an input is not specified for the app, the directory will not be created.
`/data-store/output`: A directory providing access to the output data for the app. The directory will be empty when the app starts.
`/data-store/iplant/home/`: A directory providing access to the user's home in the Data Store.
`/data-store/iplant/home/shared`: A directory providing read-only access to the community shared data in the Data Store.
The driver also creates a symbolic link `~/data-store` to the root of the volume mounts `/data-store` for convenient access in some apps, such as Jupyter Notebook.
## Limitations of the CSI Driver
While the Kubernetes Container Storage Interface (CSI) driver provides a flexible and powerful way to manage storage in Kubernetes, it does have some limitations, especially when it comes to latency-sensitive operations.
### Latency Concerns
The CSI driver introduces an abstraction layer between Kubernetes and the underlying storage system, which can add latency to storage operations. The latency becomes more significant when the underlying storage system is remote, such as the Data Store. This latency might not be noticeable for simple data retrieval tasks, but it can become significant in certain scenarios:
**High-frequency File Operations**: When dealing with operations that involve the creation, modification, or deletion of a large number of small files (e.g., tens of thousands or more), the additional overhead from the CSI layer and the data latency between the Discovery Environment and the Data Store can lead to performance bottlenecks.
**Real-time Data Transfers**: For applications that require low-latency, real-time data transfers, such as those using curl, wget, or similar tools, the latency introduced by the CSI driver might affect performance. The overhead of managing the connection between Kubernetes and the Data Store could lead to slower transfer speeds, especially for large-scale data transfers or when dealing with numerous small files.
## When Not to Use CSI
Given these limitations, it may be more effective to avoid using the CSI driver in the following scenarios:
**Real-time Data Transfer Tools**: When using tools like curl and wget for downloading or uploading data, particularly in environments where speed and low-latency are critical, it may be better to use a directory other than `/data-store` or `~/data-store`. This allows the tools to utilize local storage, which can offer low-latency than the Data Store.
**Massive File Operations**: For workloads that involve the transfer or manipulation of many thousands of files, especially if they are small, consider using a directory other than `/data-store` or `~/data-store`. This allows the workloads to utilize local storage without the additional overhead of the CSI driver.
## When to Use the CSI Driver
Despite some limitations, the Kubernetes Container Storage Interface (CSI) driver excels in many scenarios, particularly where flexibility, scalability, and ease of integration are important. Below are some examples of when using the CSI driver is highly advantageous:
### Ideal Use Cases
**Interactive Notebooks**: For environments like Jupyter Notebooks, the CSI driver is an excellent choice. Notebooks often require seamless access to persistent storage for saving and retrieving notebooks, datasets, and outputs. The CSI driver provides the necessary abstraction to manage data in the Data Store dynamically, allowing for a smooth and efficient workflow.
**Small to Medium Data Sets**: When dealing with small to medium-sized datasets, the CSI driver is well-suited for the job. These datasets can be easily managed and accessed via the volume mount paths (`/data-store` and `~/data-store`), providing users with consistent and reliable storage without the need for complex configurations.
**Persistent Storage for Stateful Applications**: Applications that require persistent storage -- such as databases, stateful services, or applications that need to maintain state between restarts -- benefit from the CSI driver’s ability to manage storage volumes effectively. The driver ensures that data remains consistent and available, even as containers are moved across nodes within the Kubernetes cluster.
## Summary
The CSI driver offers Discovery Environment Vice Apps a robust interface for accessing user and community-shared data in the Data Store. It is designed to provide flexible and user-friendly access, particularly when data access needs are moderate. It is an ideal choice for interactive environments, development workflows, and applications that need reliable and persistent access to storage without the overhead and complexity associated with direct storage management.
!!! info "Under the hood"
Installation and configuration of the driver that creates these mounts are covered in [iRODS CSI driver](https://docs.cyverse.org/deployment/05-core-services/irods-csi-driver/){target=_blank} in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/vice/csi.md){target=_blank} (last source update 2025-03-14), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/discovery-environment/vice/quick-app-launch/
---
title: "Creating a shareable quick launch button"
description: "Create saved launches, embeddable buttons, and shareable URLs that start a Discovery Environment app with preset settings."
type: Tutorial
tags:
- Quick Launch
- Discovery Environment
- Sharing
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/vice/quick-app-launch.md"
title: "CyVerse Learning Materials: docs/de/vice/quick-app-launch.md"
author: "team:cyverse"
last_modified: "2025-03-16T09:25:29-07:00"
---
# Creating a shareable quick launch button
## 1. Log into the Discovery Environment
Log into https://de.cyverse.org
If you have not yet created an account, go to the [User Portal](https://user.cyverse.org/){target=_blank} and sign up.
## 2. Select App You Want to Share
Navigate to the **Apps** tab on the left hand side menu, select the App that you want to share and click the **Details** button.

## 3. Creating Quick Launch Link
At the bottom of the Details window, find the Quick Launch Share button. Clicking it will open a speech bubble with 3 choices: **Saved Launch**, **Embedded Code** or **Shared Saved Launch URL**. Copy the one you need.


- **Saved Launch**: Saved Launch allows you to create a launch button that starts the App with a specified set of resources.
- **Embedded Code**: Creates a button that can be used in websites; The button will then redirect to the App launch page. Note: to operate the App, users will require a CyVerse account.
- **Shared Saved Launch URL**: Creates a link to the App.
## 4. Sharing with Collaborators
Here is the link created with the **Embedded Code** choice:
```
```
This can be pasted onto a webpage to create the following button:
!!! info "Under the hood"
Quick launches, which preset an app's parameters, are created and resolved through the Terrain API endpoints described in [Quick launches](https://docs.cyverse.org/api/endpoints/quick-launches/){target=_blank} in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/vice/quick-app-launch.md){target=_blank} (last source update 2025-03-16), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/discovery-environment/vice/extend-apps/
---
title: "Extending interactive apps"
description: "Copy an existing VICE app, modify an existing tool's container, or build a new Docker-based tool and interactive app in the Discovery Environment."
type: Guide
tags:
- VICE
- Docker
- App Integration
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/vice/extend_apps.md"
title: "CyVerse Learning Materials: docs/de/vice/extend_apps.md"
author: "team:cyverse"
last_modified: "2025-03-16T09:25:29-07:00"
---
# Extending interactive apps
There are several options to extend interactive apps for your specific use(s):
| Options | Hypothetical use case |
|---------|-----------------------|
| [Copy an Existing App](#copy-an-existing-app) | You want to create a Quick Launch Button which has a different set of input data, specific to a course you're teaching |
| [Modify an Existing Tool](#modify-an-existing-tool) | You need to add new packages or libraries to an existing Featured App and Tool that requires building a new Docker Container |
| [Create a New Tool & App](#create-a-new-tool) | There are no existing App types which fit your needs. You need to develop your own Docker container and integrate it from start to finish |
Depending on your specific needs, it may be easiest to copy an existing app or modify an existing tool. If you cannot find any existing Apps which suit your use case, do not hesitate to contact [CyVerse support](https://unm-carc.github.io/cyverse/getting-started/getting-help/) to ask about the availability of your favorite development environment.
??? tip "Definitions"
**Tool** - the Discovery Environment refers to Docker container templates as "Tools". The tool builder allows you to set the environment, UID, entrypoint, working directory, and computational requirements of your Docker container.
**Apps** - Applications, or "Apps", are the user interface that you build in the Discovery Environment to interact with our Tools. Apps are designed using a template builder.
Types of Tools include: `executable`, `HPC`, `OSG` (OpenScienceGrid), or `interactive`.
Apps may require input data, parameters, settings, and flags which the user selects each time they are run. `executable`, `HPC`, and `OSG` apps run non-interactively until they complete.
For `interactive` VICE apps, the App template may only include a set of input data files and folders, or nothing at all. VICE Apps use Kubernetes to orchestrate their launch.
Adding `interactive` Tools and Apps is different from adding `executable` Tools. VICE applications like Jupyter and RStudio run on open ports, enabling their User Interface (UI) in the browser.
## Copy an Existing App
[de]: https://unm-carc.github.io/cyverse/assets/de/logos/deIcon.svg
[home]: https://unm-carc.github.io/cyverse/assets/de/menu_items/homeIcon.svg
[data]: https://unm-carc.github.io/cyverse/assets/de/menu_items/dataIcon.svg
[apps]: https://unm-carc.github.io/cyverse/assets/de/menu_items/appsIcon.svg
[analysis]: https://unm-carc.github.io/cyverse/assets/de/menu_items/analysisIcon.svg
Navigating the [![][de]{width=25}](https://de.cyverse.org){target=_blank} [Discovery Environment](https://de.cyverse.org){target=_blank}:
![][apps]{width=20} **Apps** - Applications (including VICE interactive applications)
Under "Apps" you will see "Tools" and "Instant Launches" -- for this section, we are interested in "Tools".
![][analysis]{width=20} **Analyses** - Status and history of analysis jobs
Analyses will be where you can test your new App to ensure it is functioning properly.
1. If necessary, log into the [![][de]{width=25}](https://de.cyverse.org){target=_blank} [Discovery Environment](https://de.cyverse.org){target=_blank}
2. Click the [![][apps]{width=20}](https://de.cyverse.org/apps){target=_blank} [Apps](https://de.cyverse.org/apps){target=_blank} icon in the left side of the navigation view. Or use the search bar to search existing public Apps for what you're interested in.
3. When you've found the app you want, click on the three vertical dots on the right side of the selection field and select "Copy App".
4. You will be taken to the App editor, where you can now give the `Copy of App Name` a new name. You can also change the App's Description and the Tool used by the App.
6. Modify the Parameters as you see fit.
7. Save your new app. The app will be private, and is available under your "Apps Under Development" tab in [![][apps]{width=20}](https://de.cyverse.org/apps){target=_blank} [Apps](https://de.cyverse.org/apps){target=_blank}
## Modify an Existing Tool
Copying an App does not change the underlying Tool (Docker image). You may need to install some new system packages or software libraries that are too complex or take too long to compile on any of the public featured apps.
If you find that one of the existing public Apps, e.g. our Featured Apps, is useful but may be missing some key packages, you can take the next step of building a new container from our featured images.
??? tip "Why use our Featured Images?"
Some of the managed features that CyVerse provides for you are not available on publicly maintained Docker images of popular data science development environments.
For example, CyVerse adds a reverse proxy to allow RStudio to work behind our authenticated systems, and we install iRODS iCommands and other popular package managers and editors on all of our featured images.
By building a new container from our featured image set, you are assured that your new Docker image will work immediately in the Discovery Environment.
It is **strongly recommended** if you're using common IDEs like RStudio, Jupyter Lab, VS Code, Remote Desktops, or web-based applications like Shiny, Flask, or Streamlit, that you use one of our featured Docker images, or at least view our public Dockerfiles on [GitHub](https://github.com/cyverse-vice){target=_blank}, to ensure your new Container is compatible with the Discovery Environment.
### Select a Featured Base Image
Each of these featured Apps have a public GitHub repository where their Dockerfile is available. The containers are hosted on CyVerse Harbor public/private Docker container registry.
| Name | Dockerfile |
|------|------------|
| |[JupyterLab Datascience](https://github.com/cyverse-vice/jupyterlab-datascience){target=_blank} |
| | [RStudio Verse](https://github.com/cyverse-vice/rstudio-verse){target=_blank}|
| | [CloudShell](https://github.com/cyverse-vice/cli){target=_blank}|
| | [KASM Ubuntu Desktop](https://github.com/cyverse-vice/kasm-ubuntu){target=_blank} |
| | [VS Code](https://github.com/cyverse-vice/vscode){target=_blank} |
### Write your new Dockerfile
Create your own Dockerfile. We suggest using GitHub, and setting up a [GitHub Action](https://github.com/marketplace/actions/build-and-push-docker-images){target=_blank} for building your container and pushing it to a public Docker Registry.
If you select a Featured App image, it will be pulled from our public [Harbor Registry](https://harbor.cyverse.org/harbor/projects/17/repositories/){target=_blank}, which uses a different naming convention than Docker Hub containers.
For example, for the Rocker Project's featured [RStudio Verse Latest](https://github.com/cyverse-vice/rstudio-verse){target=_blank} image, the `FROM` statement would be:
```{docker}
FROM harbor.cyverse.org/vice/rstudio/verse:latest
```
You can then follow with your own `ENV` and `RUN` commands to do your own package installations.
Build your container as you normally would:
```{bash}
$ docker build -t /rstudio/verse:custom-latest .
```
??? tip "Selecting tag names"
By default Docker gives the `latest` tag to containers without a `:` and trailing tagname.
You can modify your tagname in any way you see fit, but names like `dev` and `latest` should only be used for development or activities that do not require rigorous reproducability.
It is a good practice to use versioned tag names when preparing for peer-reviewed scientific publication.
### Push your new image to a public container registry
After you have built your new image, you need to push it to a public Docker registry. We recommend [Dockerhub](https://hub.docker.com){target=_blank} or [quay.io](http://quay.io){target=_blank}. The Discovery Environment can pull any public Docker image from any public/private registry.
We currently do not allow private Docker images to be pulled, but do contact us about special private container hosting in our Harbor registry if you have sensitive data or software needs.
Alternatively, you can provide us the `Dockerfile` of your requested image and we will build the Docker image for you.
If there is no `Dockerfile` for the tool that you are interested in, contact [CyVerse support](https://unm-carc.github.io/cyverse/getting-started/getting-help/) and tell us what tool you are interesting in having us make as a VICE app.
## Create a new Tool
### Add Tool
1. If necessary, log into the [![][de]{width=25}](https://de.cyverse.org){target=_blank} [Discovery Environment](https://de.cyverse.org){target=_blank}.
2. Click the [![][apps]{width=20}](https://de.cyverse.org/apps){target=_blank} [Apps](https://de.cyverse.org/apps){target=_blank} and click on the "Manage Tools" wrench icon.
3. You'll see a list of all of the tools in the DE. Click on "More Actions" and select "Add Tool".
**Add Tool**
- `Tool name` is the name of the tool. This will appear in the DE's tool listing dialog. This is mandatory field.
- `description` is a brief description of the tool. This will appear in the DE's tool listing dialog.
- `version` is the version of the tool. This will appear in the DE's tool listing dialog. This is mandatory field.
- `Type` is the type of tool. For VICE apps, choose "interactive"; for command line applications, choose "executable".
**Container Image**
- `Image name` is the name of the image and its public registry. This is mandatory field.
- `Tag` is the image tag. If you don't specify the tag, the DE will look for the `latest` tag which is the default tag.
- `Docker Hub URL` is the URL of the image on Dockerhub.
- `Entrypoint` is the Entrypoint for your tool. Entrypoint should be present in the Docker image, and if not, you should specify it here.
- `Working Directory` is the working directory of the tool and must be filled in with the value you gathered above, e.g., `/home/jovyan/work`.
- `UID` is a number and must be filled in with the value you gathered from above. Typically `root` is `0` and default users are `1000`.
**Container Ports**
- `Ports` select the external port address that your graphic interface needs.
**Restrictions**
- `Max CPU Cores` is the number of cores for your tool, e.g., 16
- `Memory Limit` is the memory for your tool, e.g., 64 GB
- `Min Disk Space` is the minimum disk space for your tool, e.g., 200 GB
#### Required settings
##### Set the `WORKDIR`
Executable apps to not require a working directory.
The container needs to have a set working directory (for interactive apps), typically this is the home folder, e.g., `/home/jovyan` or `/home/rstudio` .
Set the `WORKDIR` in the Dockerfile; if there is no set `WORKDIR`, you can set it in the Tool Builder.
??? tip "Your Data in Your Container"
We recommend that you set the working directory of your tool to the `username` home path in a new folder called `work`, e.g., `/home/jovyan/work` or `/home/rstudio/work`.
This is because the Discovery Environment's interactive apps use a [Kubernetes container storage interface (CSI)](https://github.com/cyverse/irods-csi-driver){target=_blank} driver that connects the CyVerse Data Store to your working directory in the running container. This new mount can clobber any pre-existing files in the the container's `WORKDIR`.
##### Set the `ENTRYPOINT`
The container must have an `ENTRYPOINT` set in the Dockerfile, otherwise you must set it in the Tool itself.
1. All commonly needed dependencies are installed in the container image - you will not have `root` privileges later.
2. The default user set.
3. Disable any additional authentication (CyVerse provides CAS authentication and authorization).
4. URLs will work sanely behind a reverse proxy. If they don't, you may need to add nginx to the container.
##### Set the `PORT`
Interactive Apps rely on open ports to send display information to the browser.
Ensure the listen port for the web UI has a sane default and is set in the Dockerfile, e.g. `PORT 8888` .
You must set the port in the tool to the external port that the container is listening.
??? tip "Understanding ports in Docker containers"
For interactive containers like RStudio and JupyterLab, a conventional `docker run` execution will have the port set as `-p 8888:8888` where the port number on the left side of the `:` is the external port, and on the right the internal port. For VICE apps you need only be concerned about the external port number.
??? tip "Using a reverse proxy"
The Discovery Environment has its own authentication system, which requires us to use a reverse proxy for some containers.
Our [RStudio Server](https://github.com/cyverse-vice/rstudio-verse){target=_blank} uses `nginx` to enable reverse proxy and thus we have changed the external port to `80` instead of the Rocker-Project default `8787` port number.
??? tip "Managing ports in your new tool"
Featured VICE apps have default port options based on the app: JupyterLab apps use port `8888`, RStudio apps use port `80`, and Shiny apps use port `3838`.
It is strongly recommended you do not set the `bind to host` as `true` for your added ports when creating a new App.
## Creating a VICE app for your new tool
After your tool template has been saved, you can create an App for connecting your tool to the Discovery Environment. You can [copy an existing app](#copy-an-existing-app) and select your tool if you like an existing App's layout.
Alternatively, you can create a new app from a blank template.
??? tip "Input data"
For VICE apps, be sure to check the box "Do not pass this argument to the command line" for each option you add (for VICE, this is usually just input files and folders.
!!! info "Under the hood"
How interactive tools are deployed, proxied, and given Data Store access is described on the [VICE](https://docs.cyverse.org/deployment/06-applications/vice/){target=_blank} deployment page in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/de/vice/extend_apps.md){target=_blank} (last source update 2025-03-16), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/cloud/overview/
---
title: "Cloud services"
description: "CyVerse cloud-native services: CACAO for continuous analysis and infrastructure as code, DataWatch event notifications, and the Jetstream2 partnership."
type: Guide
tags:
- Cloud
- CACAO
- Jetstream2
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/cloud/index.md"
title: "CyVerse Learning Materials: docs/cloud/index.md"
author: "team:cyverse"
last_modified: "2022-03-23T14:01:39-07:00"
---
# Cloud services
CyVerse offers two types of cloud service:
- Continuous Analysis via { width="20" } CACAO
- Event Driven Services via { width="20"} DataWatch
CyVerse is partnered with [Jetstream 2](https://jetstream-cloud.org/){target=_blank} to develop its web interface, and to manage cloud-native applications such as Terraform, Argo Workflows, and Kubernetes.
Access to [Jetstream 2 login](https://use.jetstream-cloud.org/application){target=_blank} is managed through [ACCESS-CI](https://access-ci.org/){target=_blank}.
## Continuous Analysis
[cacao]: https://unm-carc.github.io/cyverse/assets/atmosphere/cacao.png
CyVerse built continuous-analysis frameworks for its own Atmosphere cloud (now deprecated) and for [Jetstream2](https://jetstream-cloud.org){target=_blank}.
[![][cacao]{width=30}](https://unm-carc.github.io/cyverse/cloud/cacao/) [CACAO](https://unm-carc.github.io/cyverse/cloud/cacao/) is a project enabling Continuous Analysis on Kubernetes clusters; its source is on [GitLab](https://gitlab.com/cyverse/cacao){target=_blank}.
!!! info "Under the hood"
CyVerse's cloud platforms, including CACAO and the deprecated Atmosphere, are described in [Cloud services](https://docs.cyverse.org/platform/cloud/){target=_blank} in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/cloud/index.md){target=_blank} (last source update 2022-03-23), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/cloud/cacao/
---
title: "CACAO"
description: "CACAO (Cloud Automation & Continuous Analysis Orchestration), CyVerse's infrastructure-as-code service for deploying templated workloads on Jetstream2 and other clouds."
type: Guide
tags:
- CACAO
- Cloud
- Jetstream2
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/cloud/cacao.md"
title: "CyVerse Learning Materials: docs/cloud/cacao.md"
author: "team:cyverse"
last_modified: "2025-03-14T21:51:24-07:00"
- id: cyverse-dev-docs
resource: "https://docs.cyverse.org/platform/cloud/"
title: "CyVerse Developer Documentation: Cloud services"
author: "team:cyverse"
- id: cyverse-learning-materials-overview
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/home/what_is_cyverse.md"
title: "CyVerse Learning Materials: docs/home/what_is_cyverse.md"
author: "team:cyverse"
---
# CACAO
**CACAO** (Cloud Automation & Continuous Analysis Orchestration) is CyVerse's
infrastructure-as-code service for multi-cloud deployment. Instead of building
servers by hand, you pick a template and CACAO provisions and configures the
cloud resources for you.
## What you can do with CACAO
* Provision and launch scalable resources on public research clouds and
commercial cloud providers from reusable templates.
* Stand up [Project Jupyter](https://jupyter.org){target=_blank} hubs for anywhere from a few
to thousands of users.
* Deploy open-source large language model stacks (for example Ollama or Open
WebUI) together with vector databases, backed by GPUs, in environments where
your data stays private.
## Getting access
CACAO runs primarily on [Jetstream2](https://jetstream-cloud.org){target=_blank}, CyVerse's
cloud partner. Access to Jetstream2 is managed through
[ACCESS-CI](https://access-ci.org){target=_blank}:
1. Create an account at [access-ci.org](https://access-ci.org){target=_blank}.
2. Request an ACCESS allocation that includes Jetstream2 resources.
3. Sign in to CACAO through Jetstream2 and launch a template.
The [CACAO documentation on the Jetstream2 site](https://docs.jetstream-cloud.org/ui/cacao/overview/){target=_blank}
walks through the interface, templates, and deployments. CACAO's source code
is at [gitlab.com/cyverse/cacao](https://gitlab.com/cyverse/cacao){target=_blank}.
See [AI and machine learning on CyVerse](https://unm-carc.github.io/cyverse/ai/overview/) for the GPU and
LLM templates, and [Cloud services](https://unm-carc.github.io/cyverse/cloud/overview/) for CyVerse's other cloud
offerings.
!!! info "Under the hood"
CACAO's place in CyVerse's cloud history, alongside the deprecated Atmosphere platform, is described in [Cloud services](https://docs.cyverse.org/platform/cloud/){target=_blank} in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/cloud/cacao.md){target=_blank} (last source update 2025-03-14), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/cloud/datawatch/
---
title: "DataWatch"
description: "Create DataWatch listeners that call webhooks, WebDAV, or email when files change in watched Data Store folders, using the REST API or CLI."
type: Reference
tags:
- DataWatch
- Data Store
- Events
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/cloud/datawatch.md"
title: "CyVerse Learning Materials: docs/cloud/datawatch.md"
author: "team:cyverse"
last_modified: "2025-03-14T15:53:57-07:00"
---
# DataWatch
{ width="20"} DataWatch is a service that is integrated with CyVerse DataStore, it will trigger actions based on files changes in iRODS(DataStore). To use this service, you create a DataWatch listener that will listen on one or more iRODS directories. When triggered, the listener will perform the action you defined in the listener (HTTPS webhook, email, etc.) with a list of events, each event represent changes of a single file.
> There is currently no Web UI for DataWatch, user can only access DataWatch via [REST API or CLI](#using-datawatch).
## Important Details
- Your CyVerse account must have access to the iRODS path you want to listen for file changes
- Listener will be triggered for any changes made to files under the "directory" (iRODS collection) that it watches (recursive)
- DataWatch will batch events over a period of time before sending them out, so you will receive a list of events (could be a list of one).
- minimum trigger interval for a listener is 60 seconds
## CyVerse DataStore Events
Each event has a `event` field, this indicate the type of event, the value can be:
- `created` (file is created)
- `modified` (file is updated in place)
- `moved` (file is renamed/moved)
- `removed` (file is deleted)
All event **except** `moved` type will have a `path` field, this is the iRODS path of the file.
`moved` event will have `old_path` and `new_path` field that indicate the before & after iRODS path of the move.
Examples in JSON:
```json
[
{
"event": "modified",
"path": "/iplant/xxx/XXX/mydata.json"
},
{
"event": "created",
"path": "/iplant/xxx/XXX/mydata.bin"
},
{
"event": "moved",
"old_path": "/iplant/xxx/XXX/mydata.bin",
"new_path": "/iplant/xxx/XXX/mydata_moved.bin"
},
{
"event": "removed",
"path": "/iplant/xxx/YYY/mydata.json"
}
]
```
JSON Schema
```json
{
"type": "object",
"required": [
"event"
],
"properties": {
"event": {
"type": "string",
"description": "type of event",
"enum": [
"created",
"modified",
"moved",
"removed"
]
},
"path": {
"type": "string",
"description": "path of the file that trigger the event, for created/modified/removed events, NOT for moved event",
"example": "/iplant/xxx/XXX/data.bin"
},
"old_path": {
"type": "string",
"description": "old path of the file that trigger the 'moved' event",
"example": "/iplant/xxx/XXX/old/path/data.bin"
},
"new_path": {
"type": "string",
"description": "new path of the file that trigger the 'moved' event",
"example": "/iplant/xxx/XXX/new/path/data.bin"
}
}
}
```
## Listener
There are 3 type of listener currently supported:
- webhook
- webdav
- email
## Webhook
Webhook listener will make an HTTP request to a URL you specified with list of events when triggered.
Note: URL specified in Webhook listener is expected to be reachable by datawatch.cyverse.org, if you wish to add authentication, you can set basic authentication (username & password) or additional header in listener.
### POST/PUT JSON
If the method of the webhook is `POST` or `PUT` **AND** you **didn't** specify parameters (e.g. via `-param` CLI flag), then the events will be delivered in the response body as JSON.
Example:
```
POST /api/datawatch-test HTTP/1.1
Host: example.cyverse.org
Content-Type: application/json; charset=UTF-8
Content-Length: 356
[
{
"event": "modified",
"path": "/iplant/xxx/XXX/mydata.json"
},
{
"event": "created",
"path": "/iplant/xxx/XXX/mydata.bin"
},
{
"event": "moved",
"old_path": "/iplant/xxx/XXX/mydata.bin",
"new_path": "/iplant/xxx/XXX/mydata_moved.bin"
},
{
"event": "removed",
"path": "/iplant/xxx/YYY/mydata.json"
}
]
```
### POST multipart/form-data
If one or more parameters is specified (e.g. via `-param` CLI flag), method will be override to `POST` and `Content-Type` will be set to `multipart/form-data`.
Example with parameters `foo=bar`:
```
POST /api/datawatch-test HTTP/1.1
Host: example.cyverse.org
Content-Type: multipart/form-data; boundary=92fed06b6a9e422a327e48967008ee4e9c090a42b4931a96527cdcd36edb
--92fed06b6a9e422a327e48967008ee4e9c090a42b4931a96527cdcd36edb
Content-Disposition: form-data; name="foo"
bar
--92fed06b6a9e422a327e48967008ee4e9c090a42b4931a96527cdcd36edb
Content-Disposition: form-data; name="file"
[
{
"event": "created",
"path": "/iplant/home/shared/EXAMPLE_PATH/"
}
]
--92fed06b6a9e422a327e48967008ee4e9c090a42b4931a96527cdcd36edb--
```
### GET
If you specify the webhook method to be `GET`, then the events will be included in the request `events` query parameter.
Example:
```
GET /api/datawatch-test?events=%0A%5B%0A++++%7B%0A++++++++%22event%22%3A+%22modified%22%2C%0A++++++++%22path%22%3A+%22%2Fiplant%2Fxxx%2FXXX%2Fmydata.json%22%0A++++%7D%0A%5D%0A HTTP/1.1
Host: example.cyverse.org
Content-Length: 0
```
> Note: the request line may be subjected to size limit in many web server or reverse proxy implementations. Therefore `GET` request webhook is not recommended.
## WebDAV
WebDAV listener will make an request to save a JSON file with list of events when triggered.
Note: URL specified in WebDAV listener is expected to be reachable by datawatch.cyverse.org, if you wish to add authentication, you can set basic authentication (username & password) or additional header in listener.
When WebDAV listener triggers, it will make a `PUT` request to save a json file. The json file will be named in the format like `DataWatch_2025_01_02_15_04_05.9999999_MST.json`
e.g. if the url you specified for the listener is https://example.cyverse.org/webdav/example, then the request will be made to https://example.cyverse.org/webdav/example/DataWatch_2025_01_02_15_04_05.9999999_MST.json
The content of the json file will follow the same schema as `POST/PUT JSON` webhook
## Email
The email will be sent by noreply@cyverse.org with a subject `DataWatch Notification`, the body of the email will be JSON string of the list of events.
Example:
```
From: noreply@cyverse.org
Subject: DataWatch Notification
[
{
"event": "removed",
"path": "/iplant/xxx/YYY/mydata.json"
},
{
"event": "created",
"path": "/iplant/xxx/YYY/mydata.json"
}
]
```
## Using DataWatch
There is currently no Web UI for DataWatch, user can only access DataWatch via REST API or CLI.
### CLI
#### Download CLI
- Find latest release of the CLI at https://gitlab.com/cyverse/datawatch/-/packages .
- Download the zip for your OS and CPU Architecture.
- Unzip to get the CLI executable.
- e.g. Unix/Linux command line: `unzip datawatch_cli_linux_amd64.zip`
- Consider rename the executable to `datawatch`, the instruction below assume the executable is named `datawatch`
- e.g. Unix/Linux command line: `mv datawatch_cli_linux_amd64 datawatch`
#### Using CLI
Before using commands, you need to define the following environment variable:
- `DATAWATCH_API` - base URL for API
- default: http://localhost
- Note: if this does not start with the scheme, `http://` will be prepended so if `https` is required, make sure it is included in the variable
For example:
```bash
export DATAWATCH_API=https://datawatch.cyverse.org
```
A line setting the environment variable can be added to your shell's profile or config file. For bash add it to ~/.bashrc
#### Authentication
The DataWatch API uses keycloak for authentication by default. So the CLI needs to obtain a keycloak token in order to make calls to the API. The first time you run the CLI you will be prompted to enter your keycloak username and password. They will be used to obtain a token which is then written to a file for running the CLI in the same session. The token will expire over time and so the CLI will prompt you again to enter your username and password when that happens. If for some reason you want to manually obtain a new keycloak token, run `./datawatch login`.
Login with your CyVerse credential:
```shell
$ ./datawatch login
Enter your keycloak username:
Enter your password:
```
### Examples
- Listing user (non users can only get oneself), this can be used to verify if login is successful.
```bash
./datawatch get users
```
- Here's an example of **webhook** listener. The notification will hit a `POST` endpoint with `Authorization` header.
```bash
./datawatch create listener -listenerTypeID webhook -sourceID cyverse \
-m POST -header 'Authorization: ' \
-u "https://example.com/datawatch-upload" \
-notifyInterval 60 \
/iplant/home//watched-folder
```
> Note this command takes one or more arguments (after options), each being a path that the listener will use when matching events
- Here's an example of updating URL of an existing listener (webhook).
```bash
./datawatch update listener -u https://example.com/datawatch-upload/updated-url
```
- Here's an example of a **webdav** listener. The notification will use WEBDAV to put the notification text in a file on the CyVerse data store.
```bash
./datawatch create listener -listenerTypeID webdav -sourceID cyverse -url https://data.cyverse.org/dav/iplant/home//event.json -username -password /iplant/home//test
```
- Here's an example of an **email** listener
```bash
./datawatch create listener -listenerTypeID=email -sourceID=cyverse -e "YOUR-EMAIL@example.com" "/iplant/home//analyses" "/iplant/home//test"
```
#### Command Options
- `-listenerTypeID` is required when creating listener, the value must be one of `email`, `webdav`, or `webhook`.
- `-sourceID` is required when creating listener, the value must be `cyverse`.
- `-notifyInterval` is optional. It can be used to override the default interval (in seconds) used for bundling events for a given path. Note that the value lower than 60 **will not apply**.
### Frequently Used Commands
`delete listener`, `delete user`, `get dataSources`, `get keycloakToken`, `get listeners`, `get listenerTypes`, `get users`, `update listener`, and `update user`.
### Common Error
1. Forgot to set `DATAWATCH_API` environment variable
```
Get "http://localhost/users": dial tcp 127.0.0.1:80: connect: connection refused
```
You didn't set environment variable `export DATAWATCH_API=https://datawatch.cyverse.org`
## Open API spec
DataWatch exposes a REST API that can be accessed with a keycloak token as the bearer token in the `Authorization` header, here is the [OpenAPI spec](https://gitlab.com/cyverse/datawatch/-/blob/master/docs/openapi/datawatch-openapi.yaml){target=_blank}.
### Obtaining Keycloak token
You can obtain an keycloak token by making a `GET` request to https://datawatch.cyverse.org/keycloakToken with username and password query parameter.
e.g. `curl https://datawatch.cyverse.org/keycloakToken?username=YourUsername&password=YourPassword`
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/cloud/datawatch.md){target=_blank} (last source update 2025-03-14), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/cloud/harbor/
---
title: "Harbor container registry"
description: "CyVerse's Harbor container registry at harbor.cyverse.org, which hosts the container images behind featured Discovery Environment apps."
type: Reference
tags:
- Harbor
- Containers
- Registry
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/cloud/harbor.md"
title: "CyVerse Learning Materials: docs/cloud/harbor.md"
author: "team:cyverse"
last_modified: "2025-03-14T14:47:33-07:00"
---
# Harbor container registry
CyVerse operates a [Harbor.io](https://goharbor.io){target=_blank} container registry for its featured applications.
CyVerse subscribers can request that their applications be published in the Discovery Environment, and receive hosting for their containers in the [https://harbor.cyverse.org](https://harbor.cyverse.org){target=_blank} registry.
Harbor is an open source registry that secures artifacts with policies and role-based access control, ensures images are scanned and free from vulnerabilities, and signs images as trusted.
!!! info "Under the hood"
How CyVerse deploys and operates its registry is covered on the [Harbor](https://docs.cyverse.org/deployment/04-kubernetes/harbor/){target=_blank} deployment page in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/cloud/harbor.md){target=_blank} (last source update 2025-03-14), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/ai/overview/
---
title: "AI and machine learning on CyVerse"
description: "CyVerse capabilities for AI and ML: Data Store datasets, GPU JupyterLab apps, Ollama LLMs, AI Verde, CACAO on Jetstream2, and the Data Store MCP server."
type: Guide
tags:
- AI
- Machine Learning
- GPU
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ai/index.md"
title: "CyVerse Learning Materials: docs/ai/index.md"
author: "team:cyverse"
last_modified: "2025-10-13T12:31:11-07:00"
---
# AI and machine learning on CyVerse
CyVerse provides a comprehensive suite of tools and resources for Artificial Intelligence (AI) and Machine Learning (ML) workflows, from data storage to model deployment. This section outlines the key capabilities.
## Key Capabilities
CyVerse supports various aspects of AI/ML, including:
* **Data Management:** Host and manage your AI/ML training datasets using the CyVerse Data Store. This provides a reliable and scalable location for your data.
* **Interactive Analysis:** Perform interactive model training and experimentation using popular frameworks like PyTorch and TensorFlow within JupyterLab, hosted in the CyVerse Discovery Environment. This offers a familiar and flexible coding environment.
* **Generative AI:** Run and interact with large language models (LLMs) using Ollama, also available within the Discovery Environment. This allows you to explore the capabilities of pre-trained models or fine-tune them for your specific needs.
* **Cloud Deployment:** Deploy and scale your AI/ML infrastructure on cloud platforms like Jetstream2 and Amazon Web Services (AWS) using CyVerse's Cloud Automation and Orchestration (CACAO) tool. This enables you to leverage powerful cloud resources for demanding computations.
## Supported AI/ML Approaches
CyVerse caters to a wide range of AI/ML approaches, including:
* **Machine Learning (ML):** Develop and train a broad spectrum of machine learning models, from classical algorithms to deep learning networks.
* **Natural Language Processing (NLP):** Process and analyze text data using NLP techniques and tools.
* **Predictive AI:** Build models that forecast future outcomes based on historical data.
* **Generative AI:** Utilize models that can create new content, such as text, images, or code.
## CyVerse AI VERDE
CyVerse offers a dedicated platform for AI research and education called [**AI VERDE**](https://unm-carc.github.io/cyverse/ai/verde/). This platform provides a user-friendly interface for interacting with various AI models.
* **AI VERDE Chat:** Access AI VERDE's interactive chat interface at [https://chat.cyverse.ai](https://chat.cyverse.ai){target=_blank}.
## Detailed Resource Information
CyVerse Discovery Environment features GPU instances with NVIDIA A16 GPUs for light and moderate AI applications, best for visualization and classroom teaching.
### PyTorch & TensorFlow
The Discovery Environment provides pre-configured JupyterLab instances with GPU support for [PyTorch](https://pytorch.org){target=_blank} and [TensorFlow](https://www.tensorflow.org){target=_blank}. This allows you to accelerate your model training and experimentation.
Launch in Discovery Environment:
### Cloud-Native Services with Jetstream2 and CACAO
For large-scale AI/ML workloads, CyVerse offers access to powerful cloud resources:
* **Jetstream2:** Jetstream2 provides access to NVIDIA A100 and H100 GPUs, ideal for demanding deep learning tasks. Learn more at [https://jetstream-cloud.org](https://jetstream-cloud.org){target=_blank}.
* **CyVerse CACAO:** CACAO simplifies the deployment of complex AI/ML applications on cloud platforms. It provides pre-built templates for deploying various AI/ML Web UIs, vector databases, and interactive applications, all backed by GPU resources.
### Access CyVerse Data Store from AI Agents
To access research datasets stored in CyVerse Data Store via AI Agents or other AI workloads, CyVerse offers the Data Store MCP Server. Learn how to configure your agents or workloads at [**MCP**](https://unm-carc.github.io/cyverse/ai/mcp/overview/).
**Getting Started with Cloud Resources:**
1. Create an account on [ACCESS-CI.org](https://access-ci.org){target=_blank}.
2. Request an allocation for Jetstream2 and/or CyVerse CACAO resources through the ACCESS-CI portal.
## History and Partnerships
CyVerse actively participates in several national AI initiatives and collaborates with leading research institutions.
* **AI Institutes and Synthesis Centers:** CyVerse is involved in multiple [AI Institutes and Synthesis Centers](https://cyverse.org/mlai){target=_blank} focused on advancing predictive and generative AI.
* **Partner Institutions:**
* **University of Arizona:** The Data Science Institute Data Lab offers Machine Learning Workshops leveraging CyVerse resources.
* **Iowa State University:** CyVerse supports Iowa State University projects such as [AIIRA](https://aiira.iastate.edu/){target=_blank} and [COALESCE](https://sites.google.com/view/coalescepreview/home?){target=_blank}.
* **CU Boulder:** The NSF Synthesis Center on Environmental Science, [ESIIL](https://esiil.org/working-groups){target=_blank}, utilizes CyVerse and Jetstream2 for AI/ML research (see their Open Analysis and Synthesis Infrastructure for Science (OASIS) documentation).
* **Penn State University:** The NSF Synthesis Center on Molecular and Cellular Biology, [NCEMS](https://ncems.psu.edu/working-groups/){target=_blank}, also leverages CyVerse.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ai/index.md){target=_blank} (last source update 2025-10-13), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/ai/verde/
---
title: "Accessing AI Verde"
description: "Request AI Verde access, then call its hosted large language models from Python with an API key and the chatur-chains library."
type: Tutorial
tags:
- AI Verde
- LLM
- Python
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ai/verde.md"
title: "CyVerse Learning Materials: docs/ai/verde.md"
author: "team:cyverse"
last_modified: "2025-03-14T21:51:24-07:00"
---
# Accessing AI Verde
- AI Verde is a massive LLM distribution platform details of which can be found [here](https://datascience.arizona.edu/research/tools/ai-verde){target=_blank}.
## Step 0: Get access to AI Verde
- If you are new to AI Verde drop an email to `mithunpaul@arizona.edu` detailing
- who you are (e.g. student, faculty, researcher)
- what kind of facilities of AI Verde would you like to use (e.g. Chatbot access, API access to 75+ LLM models, RAG etc)
Once an admin approves your access, you will minimally have access to a chatbot. If you asked for more programming level access, you will be provided with a unique api-token. [Here](https://aiverde-docs.cyverse.ai/api/api-token/){target=_blank} are details on how to access these keys.
Once you have the keys here are the further steps on how to access the LLM models from a python code interface.
## Step 1: Export the API key as an environment variable
```bash
export LLM_URL="https://llm-api.cyverse.ai"
export LLM_API_KEY="paste your key here"
```
## Step 2: Set up a conda environment
For example Open a linux command line window from the tools in `de.cyverse` and type
```bash
conda create --name verde python
```
```bash
conda activate verde
```
## Step 3: Install chatur chains
`pip install chatur-chains`
## Step 4: Write the Python code
```python
import os
from pathlib import Path
```
Import the API keys form the environment which you exported earlier
```python
llm_url = os.environ.get('LLM_URL')
api_key = os.environ.get('LLM_API_KEY')
```
Next use this command to import the build_llm_proxy function from chatur chains.
```python
from chains.llm_proxy import build_llm_proxy
```
Now use the code below to initiate a connection to your favorite llm.
```python
llm = build_llm_proxy(
model="anvilgpt/llama3.1:latest",
url=llm_url,
engine="OpenAI",
temperature=0.9,
api_key=api_key,
)
```
Note: The list of LLMs (for example `anvilgpt/llama3.1:latest` in the above code) you can access is dependant per user permissions. To find out what models you have access to you can use the below curl command, where `LLM_API_KEY` is the same as given to you by the AI Verde admin.
```bash
curl -s -L "https://llm-api.cyverse.ai/v1/models" -H "Authorization: Bearer $LLM_API_KEY" -H 'Content-Type: application/json'
```
Next you can start asking question to this LLM as shown below
```python
message = llm.invoke("what is the largest wildfire in arizona")
```
```python
print(message.content)
```
## Step 5: Run the code
Next save this editor file, as verde.py and run the python script from command line like this:
```bash
python verde.py
```
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ai/verde.md){target=_blank} (last source update 2025-03-14), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/ai/mcp/overview/
---
title: "AI Verde MCP servers"
description: "AI Verde Model Context Protocol (MCP) servers that let AI agents and LLM applications work with data in the CyVerse Data Store."
type: Guide
tags:
- MCP
- AI Agents
- Data Store
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ai/mcp/index.md"
title: "CyVerse Learning Materials: docs/ai/mcp/index.md"
author: "team:cyverse"
last_modified: "2025-10-13T12:31:11-07:00"
---
# AI Verde MCP servers
AI Verde **Model Context Protocol (MCP) Servers** provide a crucial bridge, allowing **Generative AI Agents** and large language models (LLMs) to securely and intelligently interact with the rich data and computational resources of CyVerse Infrastructure.
This mechanism enables AI workloads—such as RAG (Retrieval-Augmented Generation) applications, advanced chatbots, and data analysis pipelines—to access research datasets stored in the **CyVerse Data Store** and utilize other key services.
The [MCP servers contents](https://unm-carc.github.io/cyverse/ai/mcp/) page lists the configuration guides for general MCP clients, Claude Desktop, and VS Code.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ai/mcp/index.md){target=_blank} (last source update 2025-10-13), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/ai/mcp/general-configuration/
---
title: "General configuration for the AI Verde Data Store MCP server"
description: "Streamable-HTTP and SSE endpoints, HTTP authentication, and anonymous public access for the AI Verde Data Store MCP server."
type: Reference
tags:
- MCP
- Configuration
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ai/mcp/general.md"
title: "CyVerse Learning Materials: docs/ai/mcp/general.md"
author: "team:cyverse"
last_modified: "2025-10-28T15:53:57-07:00"
---
# General configuration for the AI Verde Data Store MCP server
This guide provides the necessary configuration details for using the AI Verde Data Store MCP Server.
## Streamable-HTTP
The AI Verde Data Store MCP Server supports the new Streamable-HTTP protocol. This supports OAuth 2.0 authentication, which allows you to log in interactively with your CyVerse user account while adding the AI Verde Data Store MCP Server to your MCP Client.
- Streamable-HTTP URL: https://mcp.cyverse.ai/mcp
You can access your private data at `/iplant/home//` and public community-shared data at `/iplant/home/shared/`.
## Server-Sent Events (HTTP/SSE)
The AI Verde Data Store MCP Server still supports HTTP/SSE, even though the protocol is **deprecated**. This support is maintained to ensure compatibility with many older AI Agents that can only communicate using HTTP/SSE.
- SSE URL: https://mcp.cyverse.ai/sse
You will need to set your HTTP `Authorization` header to login. Please check below for how to set the header.
You can access your private data at `/iplant/home//` and public community-shared data at `/iplant/home/shared/`.
If no user credentials are passed through the HTTP `Authorization` header, you will only be able to access public, community-shared data.
### HTTP Basic Auth
HTTP Basic Auth is a traditional but less secure way to authenticate. In this method, your CyVerse user account credentials must be passed via the HTTP `Authorization` header.
To fill the `Authorization` header, use the following format:
```
Basic .
```
You first need to generate the Base64-encoded credential string. Use your CyVerse username (e.g., `foo`) and password (e.g., `mypassword`), separated by a colon (`:`), with the following terminal command:
```bash
echo -n "foo:mypassword" | base64
```
This will generate a Base64-encoded string. Prepend this encoded text with the string `Basic ` (including the space) in the `Authorization` header while you configure your MCP Client.
## Anonymous Access
If you want to access only public community-shared data, you can use the public server.
- Streamable-HTTP URL: https://mcp-public.cyverse.ai/mcp
- SSE URL: https://mcp-public.cyverse.ai/sse
The MCP server will not require any login and grants access to `/iplant/home/shared`.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ai/mcp/general.md){target=_blank} (last source update 2025-10-28), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/ai/mcp/claude-desktop/
---
title: "Claude Desktop configuration for the AI Verde Data Store MCP server"
description: "Add the AI Verde Data Store MCP server to Claude Desktop as a custom connector for public or CyVerse-account access."
type: Guide
tags:
- MCP
- Claude Desktop
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ai/mcp/claude_desktop.md"
title: "CyVerse Learning Materials: docs/ai/mcp/claude_desktop.md"
author: "team:cyverse"
last_modified: "2025-10-28T15:53:57-07:00"
---
# Claude Desktop configuration for the AI Verde Data Store MCP server
This guide shows how to connect Claude Desktop to the AI Verde Data Store MCP server over Streamable-HTTP using the Connectors UI.
## Prerequisites
- Claude Desktop installed
- Claude Pro plan (free plan users cannot add new connectors)
- CyVerse account credentials (if not using anonymous access)
## 1. Open Connectors in Claude
1. Launch Claude Desktop.
2. Go to `File` → `Settings` → `Connectors`.
3. Click the `Add custom connector` button (Note: This button is not available to free plan users).
## 2. Add the AI Verde Data Store MCP Server
### a. Anonymous Access
Fill in the fields:
- Name: `AI-Verde Data Store Public`
- URL (anonymous public data access): `https://mcp-public.cyverse.ai/mcp`
Click the `Save` button.
### b. Full Access with CyVerse Account
Fill in the fields:
- Name: `AI-Verde Data Store`
- URL (full access): `https://mcp.cyverse.ai/mcp`
Click the `Save` button.
When adding `https://mcp.cyverse.ai/mcp`, you will be prompted for the following information to log in.
- Client ID: `mcp-client`
- Client Secret: `` **(leave it empty, as one is not required)**
## 3. Verify the Connection
Open a chat in Claude and try a simple request.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ai/mcp/claude_desktop.md){target=_blank} (last source update 2025-10-28), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/ai/mcp/vs-code/
---
title: "VS Code configuration for the AI Verde Data Store MCP server"
description: "Configure VS Code Copilot Chat to use the AI Verde Data Store MCP server with anonymous or CyVerse-account access."
type: Guide
tags:
- MCP
- VS Code
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ai/mcp/vs_code.md"
title: "CyVerse Learning Materials: docs/ai/mcp/vs_code.md"
author: "team:cyverse"
last_modified: "2025-10-28T15:42:26-07:00"
---
# VS Code configuration for the AI Verde Data Store MCP server
This document provides example configurations for setting up the remote AI Verde Data Store MCP Server for use with VS Code. Configurations are shown for using Streamable-HTTP.
## Prerequisites
- VS Code installed (latest recommended)
- Copilot Chat extension enabled (supports MCP)
- CyVerse account credentials (if not using anonymous access)
## 1. Configure MCP Server
### a. Configure MCP Server for Anonymous Access
Edit the `~/.config/Code/User/mcp.json` file.
Configure VS Code to use the AI Verde Data Store MCP Server with Streamable-HTTP. Paste the following into your `mcp.json` file.
This configuration allows access only to public data located at `/iplant/home/shared`.
```json
{
"servers": {
"public-ai-verde-datastore": {
"type": "http",
"url": "https://mcp-public.cyverse.ai/mcp"
}
}
}
```
Go to `View` → `Chat` and click the wrench icon in the chat box. Expand `MCP Server: public-ai-verde-datastore` and click `Update Tools`.
Accept adding the MCP server and opening a new popup for logging-in. Follow the instructions in the UI.
### b. Configure MCP Server with CyVerse Account
```json
{
"servers": {
"ai-verde-datastore": {
"type": "http",
"url": "https://mcp.cyverse.ai/mcp"
}
}
}
```
Go to `View` → `Chat` and click the wrench icon in the chat box. Expand `MCP Server: ai-verde-datastore` and click `Update Tools`.
Accept adding the MCP server and opening a new popup for logging-in. Follow the instructions in the UI.
VS Code will ask you to enter an OAuth client ID and secret.
- Client ID: `mcp-client`
- Client Secret: `` **(leave it empty, as one is not required)**
## 2. Verify Connection
Open Copilot Chat and try running a query, such as:
```bash
list 5 entries in /iplant/home/shared
```
## Reference
[VS Code MCP Server Documentation](https://code.visualstudio.com/docs/copilot/chat/mcp-servers){target=_blank}
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/ai/mcp/vs_code.md){target=_blank} (last source update 2025-10-28), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/education/teaching/
---
title: "Teach using CyVerse"
description: "Request a CyVerse workshop in the User Portal, choose services, add instructors, and enroll students for a class or workshop."
type: Guide
tags:
- Teaching
- Workshops
- User Portal
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/edu/index.md"
title: "CyVerse Learning Materials: docs/edu/index.md"
author: "team:cyverse"
last_modified: "2025-03-26T11:52:41-07:00"
---
# Teach using CyVerse
The [User Portal](https://user.cyverse.org){target=_blank} provides a mechanism for onboarding entire workshops or class courses, by granting your students immediate access to CyVerse platforms and services.
## For Instructors
### Step 1 Request a Workshop
Instructors start by completing the [CyVerse Resources for Training Request Form](https://user.cyverse.org/requests/8){target=_blank}.
!!! note "Avoid conflicts with Maintenance days!"
When scheduling your workshop or class, note that CyVerse conducts regular maintenance of its platforms typically on the first Tuesday of the month, with most services unavailabe during that time.
### Step 2 Select your services
In the workshop form, choose which services you want to use for your teaching. Enrollees will then have access to
those services without needing to request them.
Common platforms are: Data Store (data management), Discovery Environment (analyses), VICE (integrated environments and visualization), and [Jetstream2](https://jetstream-cloud.org/){target=_blank}(cloud computing).
### Step 3 Select your Host, Instructor(s) / Organizer(s)
You can make other CyVerse users Instructors or Organizers of your workshop; these users will have administrative rights to modify the workshop template and to add and approve student/attendee accounts.
### Step 4 Enroll your participants
In the workshop builder, you can enroll existing CyVerse users by using their first/lastname or email address or username.
You can pre-enroll new email addresses as well. We recommend using .edu, .gov, or .org email addresses.
Students can also use the URL for the workshop/class to enroll themselves. Anyone who self-enrolls subsequently must be approved by the Workshop/Class Admin or Instructor.
### Step 5 Create documentation
The CyVerse Learning Center provides templates for creating documentation for your class, e.g., Quick Starts, Guides, Tutorials, and Manuals on its [GitHub Organization](https://github.com/CyVerse-learning-materials){target=_blank}.
## For Students
In order to view the [User Portal Workshops](https://user.cyverse.org){target=_blank} page the students must first create their own CyVerse account.
[Set Up an Account](https://unm-carc.github.io/cyverse/getting-started/account/)
!!! tip "Students MUST use their institutional email addresses"
Students must use their official email addresses (.edu, .org, .gov) so that their identities can be verified by CyVerse staff.
`@gmail.com`, `@yahoo.com`, etc., email accounts will not be approved for access to the Discovery Environment Interactive Apps.
For graduate students or professionals, additional verification can be accomplished by creating and importing an [ORCID](https://orcid.org/){target=_blank}
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/edu/index.md){target=_blank} (last source update 2025-03-26), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/education/workshops/offerings/
---
title: "Workshop offerings"
description: "CyVerse's professional education series, Container Camp and FOSS, and how to run your own workshop on CyVerse resources."
type: Guide
tags:
- Workshops
- Training
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/edu/workshops/index.md"
title: "CyVerse Learning Materials: docs/edu/workshops/index.md"
author: "team:cyverse"
last_modified: "2025-03-14T14:47:33-07:00"
---
# Workshop offerings
CyVerse's popular in-person and virtual trainings are intensive, providing hands-on learning and application of cutting edge, open source technologies. Come with questions, leave with solutions.
### Professional Education Series
| Workshop | Info & Reg | What's Covered |
|----------|------------|----------------|
| Container Camp (Basic) | [Website](https://cyverse.org/cc){target=_blank} | [Syllabus](https://unm-carc.github.io/container-camp/getting-started/schedule/){target=_blank} |
| Cloud Native Camp (Advanced) | [Website](https://cyverse.org/cc){target=_blank} | [Syllabus](https://unm-carc.github.io/container-camp/getting-started/schedule/){target=_blank} |
| Foundational Open Science Skills workshops | [Website](https://cyverse.org/foss){target=_blank} | [Syllabus](https://cyverse.org/foss#schedule){target=_blank} |
### Teach your Workshops on our resources
Cyverse also helps facilitate external workshops through the [User Portal](https://user.cyverse.org/workshops){target=_blank}
Request a Workshop sign-up form be [here](https://user.cyverse.org/requests/8){target=_blank}
You can request a workshop that uses specific CyVerse platforms like the Data Store, the Discovery Environment, or CACAO and Jetstream-2.
By using the form, your students will be automatically granted access to these platforms when they enroll.
Read more detailed [instructions on setting up a class or workshop](https://unm-carc.github.io/cyverse/education/teaching/).
Questions? [Email us](mailto:info@cyverse.org)!
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/edu/workshops/index.md){target=_blank} (last source update 2025-03-14), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/education/workshops/container-camp/
---
title: "Container Camp"
description: "CyVerse Container Camp workshops on software containers for research, with links to camp materials from 2018 to 2023."
type: Reference
tags:
- Containers
- Workshops
- Docker
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/edu/workshops/container_camp.md"
title: "CyVerse Learning Materials: docs/edu/workshops/container_camp.md"
author: "team:cyverse"
last_modified: "2025-03-14T14:47:33-07:00"
---
# Container Camp
CyVerse teaches in-person and virtual workshops on the use of software containers for scientific research.
Currently, we teach both a 'beginning' and 'advanced' format for new and advanced users. Introductory camps are offered prior to advanced camps to allow users to level up or refresh their container skills.
View the schedule and sign up for the next camp here: [Container Camp Announcements](https://cyverse.org/cc){target=_blank}
| Name | Dates | Description |
|------|-------|-------------|
| [Container Camp 2023 Advanced](https://unm-carc.github.io/container-camp/getting-started/schedule/){target=_blank}| Aug 16-18 2023 | Virtual |
| [Container Camp 2023 Basics](https://unm-carc.github.io/container-camp/getting-started/schedule-basics/){target=_blank} | Mar 6-9, 2023 | Virtual |
| [Container Camp 2022](https://unm-carc.github.io/container-camp/getting-started/overview/){target=_blank} | May 12-19, 2022 | in-person & Virtual |
| [Container Camp 2021 Advanced](https://cyverse-container-camp.readthedocs-hosted.com/en/latest/){target=_blank} | Aug 2-4, 2021 | Virtual Camp |
| [Container Camp 2021 Basics](https://cyverse-container-camp.readthedocs-hosted.com/en/latest/){target=_blank} | July 26-28, 2021 | Virtual Camp |
| [Container Camp 2020](https://cyverse-container-camp-workshop-2020.readthedocs-hosted.com/en/latest/){target=_blank} | Mar 10-13, 2020 | Third camp, in person at UArizona |
| [Container Camp 2019](https://cyverse-container-camp-workshop-2019.readthedocs-hosted.com/en/latest/){target=_blank} | Mar 6-8 2019 | Second camp, in person at UArizona |
| [Container Camp 2018](https://cyverse-container-camp-workshop-2018.readthedocs-hosted.com/en/latest/){target=_blank} | Mar 7-9, 2018 | First camp, in person at UArizona |
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/edu/workshops/container_camp.md){target=_blank} (last source update 2025-03-14), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/education/workshops/foss/
---
title: "Foundational Open Science Skills"
description: "CyVerse's Foundational Open Science Skills (FOSS) workshop and an archive of its offerings from 2019 to 2024."
type: Reference
tags:
- Open Science
- Workshops
- FOSS
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/edu/workshops/foss.md"
title: "CyVerse Learning Materials: docs/edu/workshops/foss.md"
author: "team:cyverse"
last_modified: "2025-03-14T14:47:33-07:00"
---
# Foundational Open Science Skills
The most recent FOSS workshop in this archive ran virtually on Thursdays, 11:00 AM - 1:00 PM US Arizona (Mountain Standard) Time, from Sept. 5 to Nov. 21, 2024.
CyVerse's 12-week virtual workshop teaches you the principles, practices, and how-tos for doing collaborative open science using cutting-edge, open source cyberinfrastructure.
To see how our FOSS workshop can support your work, check out the workshop curriculum over the years:
| Name | Dates | Description |
|------|-------|-------------|
| [Fall FOSS 2024](https://unm-carc.github.io/foss/course/schedule/){target=_blank} | Sept. 5 - Nov. 21, 2024 | Eighth virtual workshop series
| [Spring FOSS 2024](https://github.com/CyVerse-learning-materials/foss/releases/tag/2024.spring){target=_blank} | Jan. 25 - March 14, 2024 | Seventh virtual workshop series
| [Fall FOSS 2023](https://github.com/CyVerse-learning-materials/foss/releases){target=_blank} | Sept. 7 - Nov. 2, 2023 |Sixth virtual workshop series
| [Spring FOSS 2023](https://unm-carc.github.io/foss/course/schedule/){target=_blank} | Jan 19 - Mar 30, 2023 | Fifth virtual workshop series |
| [Fall FOSS 2022](https://unm-carc.github.io/foss/course/schedule/){target=_blank} | Sept 15 - Nov 18, 2022 | Fourth virtual workshop series |
| [Fall FOSS 2021](https://cyverse-foss.readthedocs-hosted.com/en/latest/){target=_blank} | Sept 7 - Nov 18, 2021 | Third virtual workshop series |
| [Spring FOSS 2021](https://cyverse-foss.readthedocs-hosted.com/en/foss-2021-spring/){target=_blank} | Feb 9 - Apr 21, 2021 | Second virtual workshop series |
| [Summer FOSS 2020](https://cyverse-foss-2020.readthedocs-hosted.com/en/latest/){target=_blank} | July 28 - Nov 3, 2020 | First virtual workshop series |
| [Spring FOSS 2020](https://cyverse-foss.readthedocs-hosted.com/en/foss-2020-spring/){target=_blank} | Feb 17 - 21, 2020 | Second in-person workshop at UArizona |
| [Spring FOSS 2019](https://cyverse-foss.readthedocs-hosted.com/en/foss-2019-spring/){target=_blank} | Jun 3-7, 2019 | First in-person workshop at UArizona |
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/edu/workshops/foss.md){target=_blank} (last source update 2025-03-14), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/education/tutorials/science-tutorials/
---
title: "Science tutorials with CyVerse"
description: "Archive of bioinformatics, geoinformatics, and astronomy tutorials and workshops that used CyVerse, with dates and summaries."
type: Reference
tags:
- Tutorials
- Bioinformatics
- Geoinformatics
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/edu/tutorials/index.md"
title: "CyVerse Learning Materials: docs/edu/tutorials/index.md"
author: "team:cyverse"
last_modified: "2025-03-14T14:47:33-07:00"
---
# Science tutorials with CyVerse
## Bioinformatics Tutorials Using CyVerse
| Tutorial | Date | Notes |
|----------|------|-------|
|[Plant Bioinformatics Vol 3 RNA-Seq Tutorial](https://cyverse-learning-materials.github.io/pbvol3_rnaseq_tutorial/){target=_blank}| Oct 21, 2022| An end-to-end RNA-seq analysis using the Kallisto and Sleuth, emphasizing reproducibility features of the CyVerse platforms |
| [RNASeq using VICE](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736191/RNA-seq+Data+Analysis+in+CyVerse){target=_blank} | Dec 06, 2019 | Perform RNAseq differential expression analysis using Read Mapping and Transcript Assembly (RMTA) and Rstudio-DESEq2 apps |
| [Assemble a Genome Using SOAPdenovo](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736322/Assemble+a+Genome+Using+SOAPdenovo+Workflow+Tutorial){target=_blank} | Dec 19, 2019 | Commonly used procedure for de novo whole genome assembly of Illumina reads using the DE: Assemble reads, Assess assembly |
| [Cluster Orthologs and Paralogs and Assemble Custom Gene Sets](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736290/Cluster+Orthologs+and+Paralogs+and+Assemble+Custom+Gene+Sets+Workflow+Tutorial){target=_blank} | Dec 11, 2019 | Input entire protein-encoding gene or transcript repertoires from genomes of interest, and cluster homologs (orthologs and paralogs), then query clusters to assemble gene sets based on presence/absence and copy number. |
| [RNA-Seq with Kallisto and Sleuth](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736314/RNA+seq+tutorials-+Kallisto+and+Sleuth){target=_blank} | Nov 04, 2019 | Kallisto is a quick, highly-efficient software for quantifying transcript abundances in an RNA-Seq experiment. Sleuth is designed to analyze and visualize the Kallisto results in R. |
| [Kallisto-0.42.3-INDEX-QUANT-PE](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736149/Kallisto-0.42.3-INDEX-QUANT-PE+in+the+Discovery+Environment){target=_blank} | Oct 25, 2019 | Kallisto is a program for quantifying abundances of transcripts from RNA-Seq data, or more generally of target sequences using high-throughput sequencing reads. It is based on the novel idea of pseudoalignment for rapidly determining the compatibility of reads with targets, without the need for alignment. |
| [Genome Annotation with MAKER](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736243/MAKER-P+Atmosphere+Tutorial){target=_blank} | Oct 09, 2019 | This tutorial is a step-by-step guide for using SciApps to perform MAKER based annotation |
| [Evaluate and Pre-Process Sequencing Reads](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736315/Evaluate+and+Pre-Process+Sequencing+Reads+Workflow+Tutorial){target=_blank} | Jan 05, 2018 | Clean and filter Illumina reads using DE apps. |
| [Taxonomic Name Resolution Service (TNRS)](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736142/Taxonomic+Name+Resolution+Service+TNRS+Tutorial){target=_blank} | Dec 05, 2017 | Become familiar with TNRS to identify, correct, and update scientific names of plants. |
| [Submit High-throughput Sequencing Reads to NCBI Sequence Read Archive (SRA)](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736234/NCBI+Sequence+Read+Archive+SRA+Submission+Workflow+Tutorial){target=_blank} | Dec 04, 2017| The SRA is a canonical repository for sequencing data generated by high-throughput instruments. The CyVerse submission pipeline allows you to directly submit your data into an SRA-linked BioProject. |
| [Evaluate High-throughput Sequencing Reads with FastQC](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736170){target=_blank} | Aug 01, 2017 | FastQC is a popular tool for evaluating the quality of high-throughput sequencing reads such as from Illumina and PacBio. |
| [HTSeqQC](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736170/Quality+Control+for+High+Throughput+Sequence+Data+Workflow+Tutorial){target=_blank} | Aug 01, 2017 | An automated quality control analysis tool for a single and paired-end high-throughput sequencing data (HTS) generated from Illumina sequencing platforms |
| [Import data from NCBI SRA using the Discovery Environment](https://cyverse.atlassian.net/wiki/spaces/DEapps/pages/241882280/NCBI+SRA+Import+1.2){target=_blank} | Apr 04, 2017 | The NCBI Sequence Read Archive (SRA) is a repository for high-throughput sequencing reads. These are valuable data for novel analysis and reuse. You can directly import data from SRA into your Data Store using a Discovery Environment app. |
| [Discover Variants Using SAM Tools](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736284/Discover+Variants+Using+SAM+Tools+Workflow+Tutorial){target=_blank} | Oct 11, 2016 | Detect and call variants from sequence reads using Bowtie and SAM Tools. |
| [Filter, Trim, and Process High-throughput Sequencing Reads with Trimmomatic](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736205/Cleaning+up+your+reads+with+the+HTProcess+Pipeline){target=_blank} | Sep 15, 2016 | Trimmomatic is a popular application for filtering and trimming high- throughput sequencing reads. Several functions can remove populations of low quality reads, remove sequencing adaptors, and trim low-quality regions of individual reads. |
| [Characterizing Differential Expression With RNA-Seq (Without Reference Genome)](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736291/Tutorial+Characterizing+Differential+Expression+With+RNA-Seq+Without+Reference+Genome){target=_blank} | Jul 21, 2015 | Identify changes in gene expression levels between at least two sequenced transcriptome samples (18 separate tutorials) |
| [BLAST a Transcriptome](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736169/BLAST+a+Transcriptome+Workflow+Tutorial){target=_blank} | May 11, 2016 | Reduce number of transcripts and level of redundancy in an assembled transcriptome, and identify coding sequences that can be submitted to BLASTP searches. |
| [QIIME-1.9.1 for the DE](https://cyverse.atlassian.net/wiki/spaces/DEapps/pages/241881871/QIIME-1.9.1+in+Discovery+Environment){target=_blank} | Apr 12, 2016 | QIIME is an open-source bioinformatics pipeline for performing microbiome analysis from raw DNA sequencing data. |
| [mini SOAPdenovo](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736233/mini+SOAPdenovo+Tutorial){target=_blank} | Jan 04, 2016 | Gain familiarity with a commonly used procedure for de novo whole genome assembly of Illumina reads using the DE. |
| [Genome-wide Association Study (GWAS) Using a Genotyping-by-sequencing Approach](https://cyverse.atlassian.net/wiki/spaces/DEapps/pages/241882108/Genotyping+By+Sequencing+Workflow){target=_blank} | Sep 27, 2012 | Learn to identify genetic variants that are associated with a trait. |
### SciApps
| Tutorial | Date | Notes |
|----------|------|-------|
| [Association analysis with mixed models](https://cyverse.atlassian.net/wiki/spaces/Events/pages/242199602/GWAS+-+MLM){target=_blank} | Sep 18, 2013 | A genome-wide association study (or GWAS) workflow using TASSEL, EMMAX, and MLMM for mixed model analysis. |
### Atmosphere
| Tutorial | Date | Notes |
|----------|------|-------|
| [Basic Stacks](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736150/Basic+Stacks+Atmosphere+Images+Tutorial){target=_blank} | Nov 06, 2017| Use next generation sequence data produced from Reduced Representation Libraries (RRL) such as Restriction site associated (RAD) tags. |
| [fastStructure](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736220/Installing+R+packages+on+Atmosphere+Atmosphere+Images+Tutorial){target=_blank} | Oct 01, 2017| fastStructure is a fast algorithm for inferring population structure from large SNP genotype data. It is based on a variational Bayesian framework for posterior inference and is written in Python2.x. |
| [Installing R packages on Atmosphere](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736220/Installing+R+packages+on+Atmosphere+Atmosphere+Images+Tutorial){target=_blank} | Jun 23, 2016| Install R packages on Atmosphere: Launch instance, transfer files to instance, install R package, request imaging. |
| [QIIME-1.9.1 for Atmosphere](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736134/QIIME-1.9.1+Using+Atmosphere){target=_blank} | Jun 12, 2017 | QIIME is an open-source bioinformatics pipeline for performing microbiome analysis from raw DNA |sequencing data. QIIME is designed to take users from raw sequencing data generated on the Illumina or other platforms through publication quality graphics and statistics. QIIME has been applied to studies based on billions of sequences from tens of thousands of samples. |
| [rnaQUAST 1.2.0](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736153/rnaQUAST+1.2.0+using+Atmosphere){target=_blank} | May 19, 2016 | rnaQUAST is a tool for evaluating RNA-Seq assemblies using reference genome and gene data database. In addition, rnaQUAST is also capable of estimating gene database coverage by raw reads and de novo quality assessment using third-party software (STAR, TopHat, GMAP etc.). |
| [QUAST 4.0](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736216/QUAST+4.0+Using+Atmosphere){target=_blank} | May 19, 2016 | QUAST is a tool for evaluating genome assemblies by computing various metrics. |
| [rnaQUAST 1.1.0](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736254/rnaQUAST+1.1.0+using+Atmosphere){target=_blank} | May 11, 2016 | rnaQUAST is a tool for evaluating RNA-Seq assemblies using reference genome and gene data database. In addition, rnaQUAST is also capable of estimating gene database coverage by raw reads and de novo quality assessment using third-party software (STAR, TopHat, GMAP etc.). |
| [Evolinc](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736247/Evolinc+using+Atmosphere){target=_blank} | May 03, 2016| Evolinc is a two-part pipeline to identify lincRNAs from an assembled transcriptome file (.gtf output from cufflinks) and then determine the extent to which those lincRNAs are conserved in the genome and transcriptome of other species. |
| [FaST-LMM.Py v2.02](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736279/FaST-LMM.Py+v2.02+Atmosphere+Images+Tutorial){target=_blank} | Apr 19, 2016 | Introduce new users to the FaST-LMM software for GWAS analysis. |
| [KOBAS 2.0-09052014](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736176/KOBAS+2.0-09052014+Atmosphere+Images+Tutorial){target=_blank} | Apr 19, 2016| Learn how to annotate and identify using KOBAS 2.0. |
| [Validate Workflow v0.9](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736136/Validate+Workflow+v0.9+Atmosphere+Images+tutorial){target=_blank} | Apr 19, 2016| Learn to navigate the Validate Workflow. |
| [BATools 0.0.1](https://cyverse.atlassian.net/wiki/spaces/TUT/pages/258736321/BATools+0.0.1+Atmosphere+Images+Tutorial){target=_blank} | Apr 10, 2016 | Introduce new users to BATools and the BATools Wrapper Script. |
## Geoinformatics Tutorials Using CyVerse
| Tutorial | Date | Notes |
|-----------|------|-------|
| [NEON AOP Workshop](https://cyverse-2021-neon-aop-workshop.readthedocs-hosted.com/en/latest/index.html){target=_blank} | Nov 11, 2021 | Second virtual NEON AOP Workshop in collaboration with USDA-ARS and UArizona RISE Conference |
| [NEON AOP Workshop](https://cyverse-2020-neon-aop-workshop.readthedocs-hosted.com/en/latest/index.html){target=_blank} | Nov 05-07, 2020 | First virtual NEON AOP Workshop in collaboration with USDA-ARS and UArizona RISE Conference |
| [NEON-CyVerse Workshop](https://cyverse-neon-workshop-2019.readthedocs-hosted.com/en/latest/index.html){target=_blank} | Jan 09, 2019 | CyVerse Workshop taught at Battelle Inc. NEON Headquarters, Boulder CO |
| [NEON Data Science Institute](https://cyverse-neon-data-institute-2018.readthedocs-hosted.com/en/latest/index.html){target=_blank} | July 12, 2018 | NEON summer workshop taught at Battelle Inc. NEON Headquarters, Boulder CO |
## Astronomy Tutorials Using CyVerse
| Workshop | Date | Description |
|----------|------|-------------|
| [Cloud Computing Busy Week](http://bhpire.arizona.edu/2020/02/18/cloud-computing-busy-week/){target=_blank} | Feb 2-7, 2020 | a five-day busy week was conducted at the University of Arizona in order for PIRE members to work on large-scale synthetic data generation for the Event Horizon Telescope (EHT) project |
| [AstroContainers Workshop](https://astcon.github.io/2018-05-workshop/){target=_blank} | May 7-8, 2018 | designed for astronomers and astrophysicists at Steward, LSST, NOAO, and LPL |
| [MiniHackathon PIRE](https://astcon.github.io/2018-04-hackathon/){target=_blank} | Apr 11, 2018 | Docker and Jupyter for Reproducible Astronomy |
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/edu/tutorials/index.md){target=_blank} (last source update 2025-03-14), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/education/tutorials/dna-subway/
---
title: "DNA Subway"
description: "Use DNA Subway's Red, Yellow, Blue, Green, and Purple lines for genome annotation, gene-family prediction, DNA barcoding, RNA-Seq, and microbiome analysis."
type: Tutorial
tags:
- DNA Subway
- Genomics
- Education
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/edu/tutorials/dna_subway_guide.md"
title: "CyVerse Learning Materials: docs/edu/tutorials/dna_subway_guide.md"
author: "team:cyverse"
last_modified: "2025-03-14T14:47:33-07:00"
---
# DNA Subway
{width="201px" height="105px"}
### Goal
DNA Subway is an educational bioinformatics platform developed by CyVerse.
It bundles research-grade bioinformatics tools, high-performance computing, and databases into workflows with an easy-to-use interface.
"Riding" DNA Subway lines, students can predict and annotate genes in up to 150kb of DNA (Red Line), identify homologs in sequenced genomes (Yellow Line), identify species using DNA barcodes and phylogenetic trees (Blue Line), examine RNA-Seq datasets for differential transcript abundance (Green Line), and analyze metabarcoding and eDNA samples using QIIME (Purple Line).
------------------------------------------------------------------------
### Prerequisites
*In order to complete this tutorial you will need access to the following services/software*
| Prerequisite | Preparation/Notes | Link/Download |
|--------------|-------------------|---------------|
| CyVerse account | You will need a CyVerse account to complete this exercise | [Register](https://user.cyverse.org/){target=_blank} |
| DNA Subway Access | DNA Subway access is by request access | Check or request access: [CyVerse User Portal](https://user.cyverse.org/services/mine){target=_blank} |
------------------------------------------------------------------------
## DNA Subway Basics and Logging in to Subway
DNA Subway is designed to be a classroom-friendly approach to
bioinformatics. Unlike most CyVerse platforms, you can even use Subway
without registering for a CyVerse account. We do encourage you to
register however, only work from registered users can be saved. DNA
Subway uses the same open-source bioinformatics tools used by
researchers. See a [complete list of the
tools](https://dnasubway.cyverse.org/about/resources.html) provided in
the Subway pipelines.
**Some things to remember about the platform**
*Registered user and Guest user account types*
- DNA Subway access must be requested through the CyVerse user portal.
You can check if you already have access, or request access by
logging into the portal and visiting the [My
Services](https://user.cyverse.org/services/mine) page. If DNA
Subway is not listed, click on
[Available](https://user.cyverse.org/services/available){target=_blank} services to
request access.
- Guest users will not have their worked saved beyond a single DNA
Subway session. They are also disallowed from using one of the gene
predictors (FGenesH) in the genomic annotation pipeline (Red Line).
- We suggest that every student using DNA Subway obtain their own
account.
*Sample Datasets and reference data*
All Subway lines accept user data and also have sample data that can be
immediately used to create a project.
- **Red Line - Genome Annotation:** Samples of plant and animal
genomes that can be used in annotation projects
- **Yellow Line - TARGeT Search for transposons and other DNA
Sequences:** Several model plant genomes
- **Blue Line - DNA Barcoding and Phylogenetics:** Sample sequence
from plant, animal, fungal, and bacterial barcoding regions; human
mitochondrial DNA sequence
- **Green Line - RNA-Seq for differential expression:** Sample
high-throughput reads from RNA-Seq experiments
If there is a reference data set or sample sequence you would like
added, you can contact CyVerse using the DNA Subway [Contact
page](https://dnasubway.cyverse.org/feedback.html)
*Public and private projects* - DNA Subway projects are private by
default, but can be shared by making them public. Public projects are
searchable and are a great way to share data or present analysis for
grading in a classroom project.
------------------------------------------------------------------------
### *Logging into DNA Subway as a registered user*
1. Access the DNA Subway website at https://dnasubway.cyverse.org/
2. If you wish to use DNA Subway as a guest click 'Enter As Guest'
!!! note
When using DNA Subway as a guest, you will be able to work only on
the Red, Yellow, and Blue lines. Additionally, some Red Line
functionalities will be disabled. Finally, after logging out, or a
period of inactivity (\>\~ 30 min) you work will be discarded.
3. Enter your CyVerse username and CyVerse password.
### *Logging into DNA Subway as a guest user*
1. Access the DNA Subway website at https://dnasubway.cyverse.org/; click 'Enter as Guest'
------------------------------------------------------------------------
## Accessing Saved Private and Public DNA Subway Projects
DNA Subway projects are automatically saved for registered users. By
default, Subway projects are private upon creation and visible only to
you. You may make project public, in which case users will have the
ability to view those projects, but may not edit those projects.
------------------------------------------------------------------------
### Accessing Private Projects
1. Access the DNA Subway website at
2. Upon login, you will see a listing of your private projects. Access the project by clicking the project title.
3. From any DNA Subway page, you may access private projects by clicking the 'My Projects' button on the navigation menu on the left side of the page.
??? tip "Distinguishing Lines"
All projects in DNA Subway are associated with the color of their respective DNA Subway lines, and with a project ID number.
You may see the comments and species associated with the project
{width="275px" height="200px"}
??? tip "Deleting a project"
To delete a project, click the 'trash can' icon. Once deleted, all data related to that project will be lost and unrecoverable.
------------------------------------------------------------------------
### Accessing Public Projects
1. Access the DNA Subway website at ; login to Subway or enter as a guest user.
2. On the navigation menu on the left side of the screen, click 'Public Projects'
??? tip "Sorting and Search"
You can sort by project date or type, and you can search for a project by title, organism, or the name of the project owner. When searching, click the double arrow `` to search by your selected term.
------------------------------------------------------------------------
### Make a DNA Subway Project Private or Public
1. Access the DNA Subway website at ; login to Subway.
2. Access your selected project by clicking the project title.
3. Under the 'Project Information' tab, toggle the project setting to 'Public' or 'Private' as desired.
{width="238px" height="211px"}
------------------------------------------------------------------------
## Walkthrough of DNA Subway Red Line - Genome Annotation
Annotation adds features and information to a DNA sequence -- such as genes and their locations, structures, and functions.
A good introduction to annotation can be found in the paper [A beginner's guide to eukaryotic genome annotation](https://www.nature.com/nrg/journal/v13/n5/full/nrg3174.html){target=_blank}.
We'll also suggest the DNA Subway's primer on [annotation evidence](https://dnasubway.cyverse.org/project/ngs/panel/1946#){target=_blank}.
This guide contains an explanation of basic functions for this line, as wellas exercises that might be used in the classroom.
**Some things to remember about the platform**
@@ -205,156 +164,115 @@ as exercises that might be used in the classroom.
------------------------------------------------------------------------
### DNA Subway Red Line - Create an Annotation Project with Apollo
??? tip "transition away from Java"
DNA Subway is transitioning away from the original Java-based Apollo software as most popular web browsers will no longer support Java. The new Apollo is Java-free.
1. Log-in to [DNA Subway](https://dnasubway.cyverse.org/){target=+blank} -
unregistered users may 'Enter as Guest'
2. Click 'Annotate a genomic sequence.' (Red Square); select the
'Web Apollo' version
3. For 'Select Organism type' choose 'Animal' or 'Plant' and
then select the appropriate subtype.
The 'Select Organism' step will load appropriate sample
sequences and will also adjust the models used in the de novo gene
finding process.
4. For 'Select Sequence Source' select a sample sequence.
??? tip "Apollo support"
Currently, the Java-free Apollo version of Subway does not support upload of a custom DNA Sequence.
This feature is coming soon, but we will help you upload custom genomes/regions for your use in the classroom
5. (Optional) If you have a GFF file of annotated features, you may load these import these annotations from the Green Line, or from a
custom GFF file.
6. Name your project and organism (required) and give a description if desired. Click 'Continue' to proceed.
#### Example Exercise - Project Creation: Arabidopsis ChrI
In this and subsequent steps, we will annotate a 75KB section of Arabidopsis chromosome I.
1. Log-in to [DNA Subway](https://dnasubway.cyverse.org/){target=_blank} - unregistered users may 'Enter as Guest'.
2. Click 'Annotate a genomic sequence.' (Red Square); select the 'Web Apollo' version.
3. For 'Select Organism type' choose 'Plant' and then 'Dicotyledon'.
4. from 'Select a sample sequence' chose 'Arabidopsis thaliana (mouse-ear cress) chr1, 75.00 kb'.
5. Provide your project with a title, then Click 'Continue.'
??? tip "Sequence"
You can view your DNA sequence by clicking the 'Sequence' link in the 'Project Information' tab at the bottom of the page.
------------------------------------------------------------------------
### DNA Subway Red Line - Find and Mask Repetitive DNA
One you have created a Red Line Project, you may begin the process of generating and assembling predictions and evidence that can be used to annotate genes.
1. Click 'RepeatMasker'
2. When 'RepeatMasker' turns 'green' and the icon displays a 'V' (view); click 'RepeatMasker' again to view results.
{width="300px" height="200px"}
### **Example Exercise - Repeat Masking: Arabodopsis ChrI**
- **Example Sequence:** Arabidopsis thaliana (mouse-ear cress) ChrI, 75 kb
- **Tool(s):** RepeatMasker
- **Concept(s):** Non-coding DNA, sequence repeats, mobile genetic
elements (transposons)
Following the RepeatMasking steps for the Arabidopsis ChrI sample above, answer the following *discussion questions*:
1. How many hits were detected in your sample?
2. RepeatMasker reports the length of the repetitive sequences
(Length) as well as the class (Attributes).
- What is the average length of sequences identified as "simple
repeats"?
- What is the average length of sequences identified as "low
complexity"?
3. What is the total percentage of repetitive DNA in your sequence?
(Sum of the length of all repetitive sequence / sequence length
(75 kb)
??? tip "Some Useful Definitions for Repetitive Sequences"
- **Simple repeats:** 1-5bp repeats (e.g. repetitive dinucleotides 'AT' etc.)
- **Low Complexity DNA:** Poly-purine/ poly-pyrimidine stretches, or regions of extremely high AT or GC content.
- **Processed Pseudogenes, SINES, Retrotranscripts:** Non-functional RNAs present within genomic sequence.
- **Transposons (DNA, Retroviral, LINES):** Genetic elements which have the ability to be amplified and redistributed within a genome.
**Additional Investigation:** In the results table under 'Attributes' each repeat sequence is labeled "RepeatMasker#-XXX" The '#' is the
ordinal number of the hit, the XXX is the class of DNA element (e.g. "Simple_repeat" or "Low_complexity"). There are other types of
repetitive elements such as transposons and pseudogenes (e.g. Helitron and COPIA) Use online resources to learn more: ().
------------------------------------------------------------------------
### DNA Subway Red Line - Making Gene Predictions
De novo gene predictors can be run on a sample sequence to generate
predictions of gene structure and location based solely on the sequence
nucleotides.
1. Click on one or more gene prediction tools under the 'Gene
Prediction' stop. to view the results table, click the gene
predictor again once the indicator displays 'V' (view).
#### Example Exercise - Predict Genes: Arabidopsis ChrI
- **Example Sequence:** Arabidopsis thaliana (mouse-ear cress) ChrI,
75 kb
- **Tool(s):** Augustus, FGenesH, Snap, tRNA Scan
- **Concept(s):** Genomic DNA, Gene Structure, Canonical sequences
Following the gene prediction steps for the Arabidopsis ChrI sample
above, answer the following *discussion questions*:
1. Look at the 'Type' column in the gene prediction report.
Considering the Augustus results, find the 6th gene prediction
(hint: AUGUSTUS006;ID=g6) and then locate the first mention of the
term 'gene' and copy down the gene's 'start' (i.e. the starting
basepair). Note the number of times you see the term 'exon' (i.e.
number of exons predicted).
| Gene Predictor | Exon Start (bp) | Exon Stop (bp) |
| --- | --- | --- |
| Augustus | 23456 | 23684 |
| Augustus |
| Augustus |
| Augustus |
| Augustus |
2. Based on the chart, did all the gene predictors yield genes
starting at the same location? Did all the gene predictions have
the same number of exons?
3. Looking at the number of results returned by tRNA Scan, why are
they so different from results made by other predictors? Are their
places in the genome where tRNAs are more or less densely
concentrated?
**Additional Investigation:** Look for the background link at the bottom
of the DNA Subway home page and review the section entitled 'Gene
Finding'.
------------------------------------------------------------------------
### DNA Subway Red Line - Visualize predicted genes in a Genome Browser
A genome browser is an essential tool for visualization genomic data in
context. The integrated JBrowse genome browser will allow you to see the
visualized gene predictions generated so far.
1. Click 'JBrowse' and allow browser to load.
2. Zoom into a region (for example, paste the region
**1:3740638..3749063** into the location window.
!!! tip
- JBrowse will load multiple tracks of data. Since the entire genome is loaded, we recommend using the 'highlight a region' feature to help keep your place. You may also wish to record the
coordinates you are viewing as shown in the coordinates window.
- You may also adjust the settings for a particular track by clicking on the track name.
- Right-click on any gene to view additional details about that gene.
{width="400px" height="250px"}
3. Examine gene details by double-clicking on a gene to select; then
right-click to open the 'View Details' menu.
4. To view more tracks, click on 'Full-Screen View' in the
upper-left of the JBrowse window to see any additional tracks
available.
??? tip "Useful Definitions"
**Genome Browser:** A GUI (Graphical User Interface) for viewing
biological information. GBrowse (DNA Subway's Browser) is "designed to
view genomes. It displays a graphical representation of a section of a
genome, and shows the positions of genes and other functional elements.
It can be configured to show both qualitative data such as the splicing
structure of a gene, and quantitative data such as microarray expression
levels." [\[citation\]](http://gmod.org/wiki/GBrowse_FAQ)
**Track:** The individual regions of the display where information
imported into the browser. For each type (or source) of information,
there is usually an associated track.
#### Example Exercise - Visualize predicted genes: Arabidopsis ChrI
- **Example Sequence:** Arabidopsis thaliana (mouse-ear cress) ChrI,
75 kb
- **Tool(s):** Local Browser (JBrowse)
- **Concept(s):** Gene orientation/structure, transposons, chromosome
organization
Following the gene browser steps for the Arabidopsis ChrI sample above,
answer the following *discussion questions* (the locations of the genes
are given in parentheses and can be pasted into the browser):
*Considering the following genes:*
- BFN1-201 (1:3748591..3753070)
- SCAMP5-201 (1:3744556..3749035)
- STP1-201 (1:3776366..3780845)
- At1G11270.2 (1:3780041..3789000)
1. Do all the gene predictors agree with each other?
2. Which gene predictions seem to match the Ensemble genes most closely?
------------------------------------------------------------------------
### DNA Subway Red Line - Search Databases using BLAST
DNA Subway searches customized versions of UniGene and UniProt that
contain only validated plant proteins, and are free of predicted or
hypothetical proteins.
1. Click 'BLASTN'; wait until the flashing icon displays 'V' (view)
2. Click 'BLASTN' again to view the results.
3. Click 'BLASTX'; wait until the flashing icon displays 'V' (view).
4. Click 'BLASTX' again to view the results.
5. Click on 'JBrowse' and then click 'Full-screen View' in the
upper-left.
6. In the 'Available Tracks' menu, add the Blastn and Blastx
tracks.
??? tip "Useful Definitions"
**Some Useful Definitions**
- **BLAST:** Basic Local Alignment Search Tool (BLAST) is an
algorithm that search databases of biological sequence information
(e.g. DNA, RNA, or Protein sequence) and return matches. The
BLASTN program is specific to nucleotide data, and the BLASTX
algorithm works with sequence data translated into amino acid
sequences.
- **UniGene:** A database of transcript data, "each UniGene entry is
a set of transcript sequences that appear to come from the same
transcription locus (gene or expressed pseudogene), together with
information on protein similarities, gene expression, cDNA clone
reagents, and genomic location."
[\[citation\]](http://www.ncbi.nlm.nih.gov/unigene)
- **cDNA:** DNA produced by reverse transcribing mRNA using reverse
transcriptase. cDNAs are used to investigate mRNA within a
biological sample.
- **ESTs:** "Small pieces of DNA sequence (usually 200 to 500
nucleotides long) that are generated by sequencing either one or
both ends of an expressed gene. The idea is to sequence bits of
DNA that represent genes expressed in certain cells, tissues, or
organs from different organisms."
[\[citation\]](http://www.ncbi.nlm.nih.gov/About/primer/est.html)
#### Example Exercise - Search Databases using BLAST: Arabidopsis ChrI
- **Example Sequence:** Arabidopsis thaliana (mouse-ear cress) ChrI,
75 kb
- **Tool(s):** BLASTN, BLASTX, Upload Data
- **Concept(s):** RNA, cDNAs, ESTs, Biological Databases
Following the BLAST steps for the Arabidopsis ChrI sample above, answer
the following *discussion questions* (the locations of the genes are
given in parentheses and can be pasted into the browser):
1. Both BLASTN and BLASTX returns the 'Length' of your resulting
matches. Do you notice differences in the average lengths of
BLASTN and BLASTX matches? Explain.
2. Under 'Type' both BLASTN and BLASTX returns 'match' and
'match_part.' 'Match' is describing the overall length of a single
match, but individual significant matches may be fragmented, i.e.
'match_part.' Do BLASTN and BLASTX return 'match' and 'match_part'
results in different frequencies? Explain.
------------------------------------------------------------------------
### DNA Subway Red Line - Build Gene Models using Apollo
Apollo is an extension of JBrowse which allows the user to build and
edit gene models. Apollo has a number of features but in this tutorial,
we will give brief intro covering the conceptual steps.
**A. Import Blastn model to match for transcript length** Blast searches
are matched against UniGene(blastn) and UniProt(blasts). UniGen models
are derived from cDNA and ESTs (transcriptome evidence) produced by
experiment.
1. Open Apollo and zoom into a region of interest (e.g. **1:3793981..3802033**)
2. Ensure at least the following tracks are selected (on):
- Augustus (and other gene predictors: FGenesH, SNAP, etc.)
- Blastn
3. Double-click on the Blastn result, and drag this transcript into
the yellow 'User-created Annotations' section.
{width="400px" height="250px"}
**B. Select a scaffold model** Use transcriptome evidence (UniGene -
BLASTN) to select the best possible gene model for a scaffold. If no
gene model exists or significantly reflects the UniGene model, use the
UniGene model itself as a scaffold.
1. Drag a plausible model into the yellow 'User-created
Annotations' - in this case we will choose the Augustus model;
double-click the Augustus model to select the entire model and
drag into 'User-created Annotations'.
2. Adjust the Augustus model to match the 5' and 3' configuration
of the blastn model
- Delete the extraneous 5' exon (single-click to select;
right-click to delete)
- Adjust the new 5' end to match the length of the blastn-derived
transcript
- Adjust the 3' end of the Augustus-derived model (single-click
to select; use your cursor/mouse to adjust the model length) {width="420px" height="250px"}
**C. Edit model for splice sites and variants** Protein and EST data can
be used to examine possible alternative transcripts. Proteins give clues
to the actual length of the translated protein at that locus and its
reading frame. Like full length cDNAs, ESTs give valuable information on
transcript diversity. ESTs are generated by high throughput methods, and
although the data may be fragmentary, it may capture biologically
relevant information about splice variants.
1. Turn on the blastx track
2. Examine the additional evidence to consider making adjustments to
your Augustus-derived model. If you wish to make additional
isoforms of your gene:
- Double-click to select the entire Augustus-derived model
- Right-click on the model to duplicate
- Make adjustments to the model as desired
{width="420px" height="250px"}
You also have the option of adding additional [EST evidence](https://en.wikipedia.org/wiki/Expressed_sequence_tag){target=_blank}. For the Arabodopsis 75KB section, we have prepared a selection of EST data. You
will need to **close Apollo to load this data**.
1. Download the Arabidopsis ESTs for this region to your computer
from [this link](https://de.cyverse.org/dl/d/A9BED6DE-83F3-4F38-A3FE-0AA0A9AF5D53/EST_Chr1_3729956..3804955.fasta){target=_blank}
2. Click on 'Upload Data'; under "Add DNA data in FASTA format"
upload the EST file from the link in step 1.
3. Click on 'User BLASTN' to align the ESTs to this section of the
Arabidopsis genome
4. Open 'Web Apollo'. The "Blastn User" track should be loaded.
You may move this track to a convenient position on the browser
{width="420px" height="280px"}
While EST evidence is always incomplete, these sequences can help you
determine features of the gene model.
??? tip "Learn More about Gene Evidence"
- J.Craig Venter on [ESTs](http://dnaftb.dev.dnalc.org/39/av-2.html){target=_blank}
- "Dynamic Gene" [Evidence animation](http://dynamicgene.dnalc.org/evidence/evidence.html){target=_blank} (requires Flash)
**D. Determine translation start/stop sites** After making your
adjustments, you can confirm that your gene model(s) represents the
longest possible transcripts:
1. Double-click the model; right-click and select 'Set longest ORF'
**E. Compare gene model(s) with existing annotations** After making your
gene models you can compare them with existing annotations by turning on
the 'Ensemble genes' track. In this case, our work confirms the first
gene model made, but a potential isoform supported by blastx data is
likely incorrect.
{width="420px" height="150px"}
#### **Example Exercise - Build Gene Models using Apollo: Arabodopsis ChrI**
- **Example Sequence:** Arabidopsis thaliana (mouse-ear cress) ChrI,
75 kb
- **Tool(s):** Apollo
- **Concept(s):** Synthesizing multiple lines of evidence
Following the Apollo steps for the Arabidopsis ChrI sample above, answer
the following *discussion questions* (the locations of the genes are
given in parentheses and can be pasted into the browser):
1. Try annotation of the following genes and take notes on your
annotation ( right-click on the gene model, open the 'Information
Editor' and scroll down to the comments section to enter
comments). How do your annotations compare with the Ensembl
annotations?
*Genes to try:*
- AT1G11270.2 (1:3781511..3790520)
- STP1-201 (1:3776261..3785270)
- T28P6.11-201 (1:3762877..3764678)
------------------------------------------------------------------------
## Walkthrough of DNA Subway Yellow Line - Sequence Detection
Genome prospecting uses a query sequence (DNA or protein of up to 10,000
base pairs/amino acids) to find related sequences in specific genomes or
in a database. A major purpose of genome prospecting is to identify
members of gene or transposon families. DNA Subway uses the TARGeT
workflow, which integrates BLAST searches, multiple sequence alignments,
and tree-drawing utilities. Yellow line uses TARGeT (Tree Analysis of
Related Genes and Transposons) uses either a DNA or amino acid 'seed'
query to: (i) automatically identify and retrieve gene family homologs
from a genomic database, (ii) characterize gene structure and (iii)
perform phylogenetic analysis. Due to its high speed, TARGeT is also
able to characterize very large gene families, including transposable
elements (TEs).
[\[citation\]](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2699529/)
**Some things to remember about the platform**
- Yellow Line will return sequences that would normally be excluded
from a BLAST search of a genome (e.g. repetitive sequences,
transposons).
- Yellow Line is implemented only for plant genomes
------------------------------------------------------------------------
### DNA Subway Yellow Line - Create a Yellow Line Project
1. Log-in to [DNA Subway](https://dnasubway.cyverse.org/){target=_blank} -
unregistered users may 'Enter as Guest'
2. Click 'Prospect Genomes using TARGeT' (Yellow Square)
3. Select a sample sequence, or paste in a sequence to search for.
!!! note
DNA Subway Yellow Line is only implemented to search a limited set of plant genomes.
4. Provide your project with a title, then Click 'Continue'
#### Example Exercise - Project Creation: mPing Mite element to search plant genomes for an active transposon
The [mPing MITE element](https://www.nature.com/nature/journal/v421/n6919/full/nature01214.html){target=_blank}
is an example of an active transposon in rice.
[Transposons](http://www.dnaftb.org/32/animation.html){target=_blank} are a major class
of DNA elements that impact the function of the genome.
1. Create a Yellow Line project following the steps above and using the mPing Mite Element (Oryza sativa/Rice)
### DNA Subway Yellow Line - Search Plant Genomes with TARGeT
1. Click and select the genome(s) you wish to search and the click;
'Run' to search those genomes.
2. Click the 'Alignment Viewer' button to view the results of the
search as a multiple alignment.
3. Click the 'Tree Viewer' button to view a tree that will group
results by similarity.
??? tip "Viewer Tips"
**Alignment Viewer** Generates an alignment of all search results
{width="420px" height="150px"}
**Tree Viewer** Displays the results of sequence matches as a tree, grouped by sequence similarity
[yellow_tree](https://unm-carc.github.io/cyverse/education/tutorials/assets/dna_subway/yellow_tree.png){width="420px" height="300px"}
??? tip "Useful Definitions"
- **Transposons (DNA, Retroviral, LINES):** Genetic elements which have the ability to be amplified and redistributed within a genome.
- **Non-autonomous transposons:** Transposons which lack an active transposase gene, thus requiring help from another transposon to move.
- **Autonomous transposons:** Transposons which have a functional transposase and can move within the genome.
#### Example Exercise - Search Plant Genomes: mPing Mite element
1. After loading the mPing Mite Element as the query, search the
Oryza Sativa genome, and examine the results in the Alignment and
Tree Viewers.
2. Repeat this analysis with a new project using the Ping transposase
gene and the Ping Transposase protein.
------------------------------------------------------------------------
## Walkthrough of DNA Subway Blue Line - DNA Barcoding and Phylogenetics
You can analyze relationships between DNA sequences by comparing them to
a set of sequences you have compiled yourself, or by comparing your
sequences to other that have been published in database such as GenBank
(National Center for Biotechnology Information). Generating a
phylogenetic tree from DNA sequences derived from related species can
also allow you to draw inferences about how these species may be
related. By sequencing variable sections of DNA (barcode regions) you
can also use the Blue Line to help you identify an unknown species, or
publish a DNA barcode for a species you have identified, but which is
not represented in published databases like GenBank.
**Some things to remember about the platform**
- Wet lab protocols and other resources are available at
- The DNA Barcoding 101 site also contains information on low-cost
sequencing for U.S.-based educators.
------------------------------------------------------------------------
!!! warning "Sample Data"
**How to use provided sample data**
In this guide, we will use a mosquito dataset that includes DNA
sequences isolated from mosquito larvae collected from Virginia's
Shenandoah Valley (*"Mosquito dataset"*). There is a complete
two-hour classroom bioinformatics lab with detailed instructions for
instructors and students on QUBES hub
[here](https://qubeshub.org/qubesresources/publications/165/2){target=_blank}. Where
appropriate, a note (in this orange colored background) in the
instructions will indicate which options to select to make use of this
provided dataset.
**Sample data citation**: Williams, J., Enke, R. A., Hyman, O.,
Lescak, E., Donovan, S. S., Tapprich, W., Ryder, E. F. (2018). Using
DNA Subway to Analyze Sequence Relationships. (Version 2.0). QUBES
Educational Resources.
[doi:10.25334/Q4J111](http://dx.doi.org/10.25334/Q4J111){target=_blank}
**Video Course**
Here is a video series on analyzing data with DNA Subway using the
above mosquito dataset and lesson:
!!! tip
See a Course Source paper with protocols and recommendations for
implementing a Barcoding CURE (course-based undergraduate research
experience): [CURE-all: Large Scale Implementation of Authentic DNA Barcoding Research into First-Year Biology Curriculum](https://www.coursesource.org/courses/cure-all-large-scale-implementation-of-authentic-dna-barcoding-research-into-first-year){target=_blank}.
### DNA Subway Blue Line - Create a Barcoding Project
1. Log-in to [DNA Subway](https://dnasubway.cyverse.org/){target=_blank} - unregistered users may 'Enter as Guest'.
!!! note
Only registered users submitting novel, high-quality sequences
will be able to submit sequence to GenBank
2. Choose a project type:
**- Phylogenetics**: build phylogenetic trees from any DNA, protein, or mtDNA sequence)
**- Barcoding**: DNA Barcoding for plants (rbcL), animals (COI), bacteria (16S), and fungi (ITS).
!!! warning "Sample Data"
*"Mosquito"* dataset: Select **COI**.
3\. Under 'Select Sequence Source' select a sequence by uploading
either a FASTA file or AB1 Sanger sequencing tracefile; pasting in
a sequence in FASTA format, or selecting and importing a trace
file from DNALC. If you do not have a file, you may select any of
the available sample sequences.
!!! warning "Sample Data"
*"Mosquito"* dataset:
From **Select a set of sample sequences** select **Intro to Barcoding Bioinformatics: Mosquitoes**.
4\. Name your project, and give a description if desired; click 'Continue.'
------------------------------------------------------------------------
### DNA Subway Blue Line - View and Clean Barcoding Sequence Data
**A. View Sequencing Trace File**
If you provided AB1 trace files, or imported files from DNALC, you will
be able to view the sequence electropherogram.
1. Click 'Sequence Viewer' to show a list of your sequences.
2. Click on a sequence name to show the sequences' trace file.
{width="400px" height="200px"}
**B. Trim sequence, reverse complement and pair**
By default, DNA Subway assumes that all reads are in the forward
orientation, and displays an 'F' to the right of the sequence. If any
sequence is not in that orientation, click the "F" to reverse compliment
the sequence. The sequence will display an "R" to indicate the change.
1. Click 'Sequence Trimmer.'
2. Click 'Sequence Trimmer' again to examine to changes made in the sequence
3. Click 'Pair Builder.'
4. Select the check boxes next to the sequences that represent bidirectional reads of the same sequence set. Alternatively Select the 'Auto Pair' function and verify the pairs generated.
!!! warning "Sample Data"
*"Mosquito"* dataset:
Click **Try Auto Pairing**. One pair of horsefly sequences and 4
pairs of mosquito sequences will be created. Finally, click
**Save**.
5. As necessary, Reverse Compliment sequences that were sequenced in
the reverse orientation by clicking the 'F' next to the sequence
name. The 'F' will become an 'R' to indicate the sequence has been
reverse complimented.
6. Click **Save** to save the created pairs.
**C. Build a consensus sequence** This step remove poor quality areas at
the 5' and/or 3' ends of the consensus sequence.
1. Click on "Trim Consensus." Once the job is ready to view, click
"Trim Consensus" again to view the results. Scroll left and
right in the consensus editor window to identify what string of
nucleotides from the consensus sequence you want to trim.
2. Click on the last consensus sequence nucleotide that you want to
trim. A red line will indicate what nucleotides will be removed
from the consensus sequences.
3. Click **Trim**. A new "Consensus
Editor" window will pop up displaying the trimmed sequences.
!!! warning "Sample Data"
*"Mosquito"* dataset: All of the sequences in this dataset benefit from trimming. Follow the steps above to trim sequences. We recommending trimming at the first and last "grey" (lower quality) nucleotide on the right and left ends.
------------------------------------------------------------------------
### DNA Subway Blue Line - Find Matches with BLAST
DNA Subway Blue Line will search a local copy of a BLAST databases to
check for published matches in GenBank.
!!! tip
At the end of the BLAST results page, you can see the latest update to the DNA Subway BLAST database.
1. Click 'BLASTN' then click the 'BLAST' link to BLAST the
sequence of interest. When the search is completed a 'View' link
will appear.
2. Examine the BLAST matches for candidate identification. Clicking
the species name given in the BLAST hit will also give additional
information/photos of the listed species.
3. If desired, select the check box next to any hit, and click
**Add BLAST hits to project** to
add selected sequences to your project.
{width="400px" height="200px"}
!!! warning "Sample Data"
*"Mosquito"* dataset: We recommend performing a BLASTN search for all samples and saving the top 2 matches to your project for additional analysis (as in Step 3).
------------------------------------------------------------------------
### DNA Subway Blue Line - Add Reference Data
Depending on the project type you have created, you will have access to
additional sequence data that may be of interest. For example, if you
are doing a DNA barcoding project using the rbcL gene, samples of rbcL
sequence from major plant groups (Angiosperms, Gymnosperms, etc.) will
be provided. Choose any data set to add it to your analysis; you will be
able to include or exclude individual sequences within the set in the
next step.
1. Click 'Reference Data.'
2. Select sequences of your choice.
3. Click **Add ref data** to add
the data to your project.
!!! warning "Sample Data"
*"Mosquito"* dataset: Select **Common insects** and then click **Add ref data**.
------------------------------------------------------------------------
### DNA Subway Blue Line - Build a Multiple Sequence Alignment and Phylogenetic Tree
**A. Build a multiple sequence alignment and phylogenetic tree**
1. Click 'Select Data.'
2. Select any and all sequences you wish to add to your tree.
!!! warning "Sample Data"
*"Mosquito"* dataset: We suggest first adding your \"user data\" and building an alignment and tree. You can return to this step later to build additional trees. Once Selected, click **Save Selections**. Follow the rest of the steps in this section and section B to create your tree.
3. Click **Save Selections** to select data
4. Click 'MUSCLE.' to run the MUSCLE program.
5. Click 'MUSCLE' again to open the sequence alignment window.
{width="400px" height="200px"}
6. Examine the alignment and then select the **Trim Alignment** button in the upper-left of the Alignment viewer'.
**B. Build phylogenetic tree**
1. Click 'PHYLIP NJ' and then click again to examine a
neighbor-joining tree
{width="400px" height="200px"}
2. Click 'PHYLIP ML' and then click again to examine a
maximum-likelihood tree
{width="400px" height="200px"}
!!! warning "Sample Data"
*"Mosquito"* dataset: We suggest setting "horsefly" as outgroup for both trees.
------------------------------------------------------------------------
## Walkthrough of DNA Subway Green Line: Kallisto/Sleuth RNA-Seq
The Green Line runs within CyVerse DNA Subway and leverages powerful
computing and data storage infrastructure and uses the supercomputer
cluster to provide a high performance analytical platform with a simple
user interface suitable for both teaching and research. is a quick,
highly-efficient software for quantifying transcript abundances in an
RNA-Seq experiment. Even on a typical laptop, Kallisto can quantify 30
million reads in less than 3 minutes. Integrated into CyVerse, you can
take advantage of CyVerse DNA Subway to process your reads, do the
Kallisto quantification, and analyze reads with the Kallisto companion
software in an R-Shiny app.
**Some things to remember about the platform**
- You must be a registered CyVerse user to use Green Line.
- The Green Line was designed to make RNA-Seq data analysis
"simple". However, we ask that users thoughtfully decide what
"jobs" they want to submit. **Each user is limited to a maximum of
4 concurrent jobs running on Green Line**.
- A single Green Line project may take a week to process since HPC
computing is subject to queues which hundreds of other jobs may be
staging for. Additionally these systems undergo regular maintenance
and are subject to periodic disruption.
!!! note
**New, faster Green Line**
Green Line is now running on [Jestream Cloud](https://jetstream-cloud.org/){target=_blank}. This should greatly reduce queue times (The entire running time for this tutorial is about 60
minutes). We have designed Green Line for a lower number of
concurrent users (<50), and still recommend teaching using jobs
you have made public, and only running the entire workflow when
you are working with novel data. Please let us know about your
experience: [send feedback](https://dnasubway.cyverse.org/feedback.html){target=_blank}.
!!! danger "Important: Discontinued Support for Tuxedo Workflow"
The Tuxedo workflow previously implemented for the Green Line will
has been removed in **June 2019**. Data and previously analyzed results will still be available on the CyVerse Data Store, however it is not possible to execute new analyses which include Tuxuedo.
------------------------------------------------------------------------
!!! warning "Sample Data"
**How to use provided sample data**
In this guide, we will use an RNA-Seq dataset (*"Zika infected
hNPCs"*). This experiment compared human neuroprogenetor cells
(hNPCs) infected with the Zika virus to non-infected hNPCs. You can
read more about the experimental conditions and methods in this [reference](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0175744){target=_blank}.
Where appropriate, a note (in this orange colored background) in the
instructions will indicate which options to select to make use of this
provided dataset.
**Sample data citation**: Yi L, Pimentel H, Pachter L (2017) Zika
infection of neural progenitor cells perturbs transcription in
neurodevelopmental pathways. PLOS ONE 12(4): e0175744. [reference](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0175744){target=_blank}.
**Video Course**
Here is a video series on analyzing data with DNA Subway using the above Zika dataset and lesson:
### DNA Subway Green Line: Kallisto/Sleuth - Create an RNA-Seq Project to Examine Differential Abundance
**A. Create a project in Subway**
1. Log-in to - unregistered users may NOT use Green Line.
2. Click on the Green "Next Generation Sequencing" square to start
a Green Line project.
3. For 'Select Project Type' select either "Single End Reads" or
"Paired End Reads".
!!! warning "Sample Data"
*"Zika infected hNPCs"* dataset: Select **Paired End Reads**
4\. For 'Select an Organism' select a species and genome build.
!!! warning "Sample Data"
*"Zika infected hNPCs"* dataset: Select **Homo sapiens - Ensembl 78 GrCh38**
5\. Enter a project title, and description; click 'Continue'.
!!! tip
If you don't see a desired species/genome [contact us](https://dnasubway.cyverse.org/feedback.html){target=_blank} to have it added.
**B. Upload Read Data to CyVerse Data Store** The sequence read files
used in these experiments are too large to upload using the Subway
internet interface. You must upload your files (either .fastq or
.fastq.gz) directly to the CyVerse Data Store.
1. Upload your reads to the CyVerse Data Store using Cyberduck. See instructions: [Data Store Guide](https://unm-carc.github.io/cyverse/data-store/sftp/cyberduck/).
!!! note
This step is not directly connected with DNA Subway. You can use any data uploaded to the CyVerse Data Store.
!!! warning "Data Limit"
There is a limit of 6GB per file for samples on Green Line. For larger file sizes, you may wish to use the Kallisto tools in the CyVerse Discovery Environment. See the for more information.
------------------------------------------------------------------------
### DNA Subway Green Line: Kallisto/Sleuth - Manage Data and Check Quality with FASTQC
**A. Select and pair files**
1. Click on the "Manage Data" step: this opens a Data store window
that says "Select your FASTQ files from the Data Store" (if you
are not logged in to CyVerse, it will ask you to do so).
2. Click on the folder that matches your CyVerse username and
Navigate to the folder where your sequencing files are located.
!!! warning "Sample Data"
*"Zika infected hNPCs"* dataset: Select **Sample Data**.
3. Select the sequencing files you want to analyze (either .fastq or .fastq.gz format).
!!! warning "Sample Data"
*"Zika infected hNPCs"* dataset: You will be presented with the following 8 files; **check-select all of the files** and click the **+ Add files** button:
- SRR3191542_1.fastq.gz
- SRR3191542_2.fastq.gz
- SRR3191543_1.fastq.gz
- SRR3191543_2.fastq.gz
- SRR3191544_1.fastq.gz
- SRR3191544_2.fastq.gz
- SRR3191545_1.fastq.gz
- SRR3191545_2.fastq.gz
The SRR3191542 and SRR3191543 files are 2 replicates (paired-end) of the uninfected cells and the SRR3191544 and SRR3191545 file are from the Zika infected cells.
4. If working with paired-end reads, click the **Pair Mode OFF** button to toggle to on; check each pair of sequencing files to pair them.
!!! warning "Sample Data"
*"Zika infected hNPCs"* dataset: Right reads end in "_1" and left reads end in "_2". **Click the** **Pair Mode OFF** **button** to turn pairing on, and **check-select each of the paired samples** (e.g. SRR3191543_1.fastq.gz and SRR3191543_2.fastq.gz).
**B. Check sequencing quality with FastQC**
It is important to only work with high quality data. is a popular tool
for determining sequencing quality.
!!! tip
This step takes place in the same **Manage data** window as the steps above.
1. Once files have been loaded, in the 'Manage Data' window, click the 'Run' link in the 'QC' column to run FastQC.
!!! note
There is a limit of 4 concurrent jobs. These jobs should take less
than 20 minutes to complete (depending on file size) and you may need
to let several jobs finish before proceeding. If you have previously
processed reads for quality, you can skip the FastQC step.
2\. One the jobs are complete, click the 'View' link to view the results.
!!! tip
You can see a description and explanation of the FastQC report on the CyVerse Learning Center and a more detailed set of explanations on the website.
------------------------------------------------------------------------
### DNA Subway Green Line: Kallisto/Sleuth - Trim and Filter Reads with FastX Toolkit
Raw reads are first "quality trimmed" (remove poor quality bases off
the end(s) of a read) and then are "quality filtered" (filter out
entire poor quality reads) prior to aligning to the transcriptome. After
trimming and filtering, FastQC is run on the trimmed/filtered files.
1. Click "FastX ToolKit" to open the FastX Toolkit panel for all your data.
2. For each file, under 'Basic', Click 'Run' to filter the reads using default parameters or click 'Advanced' to run with desired parameters; repeat this process for all the FASTQ files in your dataset.
!!! warning "Sample Data"
*"Zika infected hNPCs"* dataset: The quality of the reads in this dataset is relatively good. You can **skip the FastX Toolkit step for this dataset**.
!!! tip
The 'Basic' setting for FastX Toolkit uses the same settings as
the defaults in the 'Advanced' run:
- **quality_trimmer: minimum quality**: 20
- **quality_trimmer: minimum trimmed read length**: 20
- **quality_filter: minimum quality**: 20
- **quality_filter: minimum quality**: 50
3. Once the job completes, click the 'View' link to view a generated FastQC report.
4. Since you may trim reads multiple times to achieve the desired quality of data record the job IDs (e.g. fx####) that you wish to use in the subsequent steps.
------------------------------------------------------------------------
### DNA Subway Green Line: Kallisto/Sleuth - Quantify reads with Kallisto
Kallisto uses a 'hash-based' pseudo alignment to deliver extremely fast
matching of RNA-Seq reads against the transcriptome index (which was
selected when you created your Green Line project). A Kallisto analysis
must be run for each mapping of RNA-Seq reads to the index. In this
tutorial, we have 12 fastQ files (6 pairs), so you will need to launch 6
Kallisto analyses.
??? tip "The Science Behind Kallisto"
You can find a detailed video series on the science behind the Kallisto software and pseudoalignment: [YouTube](https://www.youtube.com/playlist?list=PL-0S9LiUi0vhjynujVZw34RKmUo6vPmVd){target=_blank}.
1. Click the "Quantification" step and enter a sample and condition name for each of your samples. You will typically have several replicates (at least 3 minimum) for each sample. For your condition, our implementation of the Kallisto/Sleuth workflow supports **two conditions**.
!!! danger "Warning"
When naming your samples and conditions, avoid spaces and special
characters (e.g. !#\$%\^&/, etc.). Also be sure to be consistent with spelling.
!!! warning "Sample Data"
*"Zika infected hNPCs"* dataset:
We suggest the following names for this dataset:
| Left/Right Pair | Sample name | Condition |
| --- | --- | --- |
| SRR3191542_1.fastq.gz SRR3191542_2.fastq.gz | Mock1-1 | Mock |
| SRR3191543_1.fastq.gz SRR3191543_1.fastq.gz | Mock2-1 | Mock |
| SRR3191544_1.fastq.gz SRR3191544_2.fastq.gz | ZIKV1-1 | Zika |
| SRR3191545_1.fastq.gz SRR3191545_2.fastq.gz | ZIKV2-1 | Zika |
2. After naming the samples and conditions, click the **Submit** button to submit a job. Typically, within \~1 minute you will be provided with a job number. The job will be entered into the queue at the TACC Stampede supercomputing system. You can come back and click the Quantification stop to see the status of the job. The indication for the quantification stop will show "R" (running) while the job is running.
!!! warning "Sample Data"
*"Zika infected hNPCs"* dataset: Under parameters **uncheck** the *Build pseudo-bam files* option.
!!! tip
You can select some of the advanced options for your Kallisto job by clicking the "Parameters" link in the Quantification stop. See more about these advanced parameters in the [Kallisto manual](https://pachterlab.github.io/kallisto/manual){target=_blank}.
------------------------------------------------------------------------
### DNA Subway Green Line: Kallisto/Sleuth- Visualize data using IGV
In the "View Results" steps you have access to alignment
visualizations, data download, and interactive visualization of your
differential expression results.
1. Click the "View results" step and choose one of the following
options:
**IVG - Integrated Genome Viewer**
!!! tip
IGV visualization will only be possible if you have built pseudo-bam
files in the Kallisto step.
Click the icon in the "IGV" column to view a visualization of
your reads pseudoaligned to the reference transcriptome. You will need
to click the **Make it public**
button (and possibly be re-directed to the CyVerse Discovery
Environment). After making the data "public" which allows DNA Subway
to access your files on the CyVerse Data Store, you must also select a
memory size to launch this Java application. If you are not sure of
which value to select, use the default 750MB option.
!!! warning
Using IGV requires Java software. Java is increasingly unsupported for
security reasons on the internet.
!!! tip "Java Help"
Java must be available and enabled in your Internet browser to use the
IGV function. Java frequently is the source of security
vulnerabilities and so its not uncommon to experience configuration
issues due to safety. Follow the tips below to configure Java for your
computer. Alternatively, you can use the Download link (see
instructions in the section below) to download your data (you will
need the .bam and .bam.bai files) and download and install yourself.
*Internet Browser*
We highly recommend using Firefox as your browser for DNA Subway.
- Verify your Java availability for your browser: [Java test](https://www.java.com/en/download/installed.jsp){target=_blank}
- Java must be [enabled](https://java.com/en/download/help/enable_browser.xml){target=_blank} in your browser
*Java Configuration*
- Open the Java control panel on your computer. (On Mac, open System
Preferences > Java. On PC, open Control Panel > Programs >
Java.)
- Click the Security tab and check "Enable Java in the browser"
and set the security level for applications to "high". Add
"" and ""
to the "Exception Site List" in the Java Security tab.
**Download Data - Abundance**
Click the folder icon to be redirected to the CyVerse Discovery
Environment (you may be required to log in). You will be directed
to all outputs from you Kallisto analysis. You may preview them in
the Discovery Environment or use the path listed to download the
files using Cyberduck (see [Data Store Guide](https://unm-carc.github.io/cyverse/data-store/sftp/cyberduck/)). A tab-separated file of abundances
for each sequence pair is available at the download link.
------------------------------------------------------------------------
### DNA Subway Green Line: Kallisto/Sleuth- Visualize data using Sleuth
**Differential analysis - Shiny App**
Click the "Sleuth R Shiny" link to launch an interactive window which contains data and graphics from your analysis.
**R Shiny App Walkthrough**
The R Shiny App allows you to explore your differential expression
results as generated by the . We will cover highlights to for each
menu in the app.
??? tip "Data Transfer Timings"
It can take a few minutes for data to be transferred to the R
Shiny server after the quantification step completes. If R Shiny
does not load, try again in a few minutes. If you still have an
issue, use the link and include your project number in the
feedback form.
**Results Menu**
{width="800px" height="400px"}
This menu is an interactive table of your results. You can choose
which columns to display in the table using the checkboxes on the
left of the screen. Several important values selected by default
include:
- **Target_id**: This is the name of the transcript (gene) from the selected reference transcriptome.
- **qval**: This is a corrected (for multiple testing) p-value indicating the significance test of differential abundance. Lower numbers indicate greater significance.
- **b**: This is an estimate of the fold change between the conditions
- **ext_gene**: If available, these are gene names pulled from Ensemble
!!! tip
Click the **Download** button to download these results.
**Bootstrap**
{width="800px" height="400px"}
This menu will display a box plot that indicates the difference in
expression between conditions. The box plots themselves indicate
variation between replicates as estimated by bootstrap sampling of
the reads. A dropbox enables you to select any transcript.
Clicking the "Show genes" will load alternative gene names if
available.
!!! tip
Right-click a graph to download this and other images.
**PCA**
{width="800px" height="400px"}
This graph displays principle components of each of the
conditions/replicates. In general replicates of the same condition
should cluster closely together.
**Volcano Plot**
{width="800px" height="400px"}
This scatter plot displays all transcripts colored by significance
of differential abundance. You may also use menu on the left of
the screen to highlight specific genes/transcripts or previously
set filters from the results menu.
**Loadings**
{width="800px" height="400px"}
This barplot indicates which genes/transcripts explain most of the variance computed in the principle components analysis.
**Heatmap**
{width="800px" height="400px"}
This heatmap gives a measure of the similarity between the possible comparison of the samples and their replicates.
------------------------------------------------------------------------
**Summary**: Together, Kallisto and Sleuth are quick, powerful ways to
analyze RNA-Seq data.
------------------------------------------------------------------------
## Walkthrough of DNA Subway Purple Line (beta testing documentation)
!!! danger "BETA RELEASE"
The Purple line is in beta release. Please send feedback to [DNALC Admin](mailto:dnalcadmin@cshl.edu).
The Purple Line provides the capability for analysis of microbiome and
eDNA (environmental DNA) by implementing a simplified version of the [QIIME 2](https://qiime2.org/){target=_blank}
(pronounced "chime two") workflow. Using the Purple Line, you can
analyze uploaded high throughput sequencing reads to identify species in
microbial or environmental DNA samples.
Metabarcoding uses high-throughput sequencing to analyze hundreds of
thousands of DNA barcodes from complex mixtures of DNA. In a typical
experiment, DNA is isolated from sterile swabs or material taken from
different environmental locations or conditions. PCR is used to amplify
a variable region, such as COI, or 12S or 16S ribosomal RNA genes, and
sequence reads identify the variety and abundance of species from
different samples. The analysis requires specialized software, such as
QIIME 2.
The Purple Line integrates sequence data and metadata imported from
CyVerse's Data Store, demultiplexing of samples, quality control, and
taxonomic identification and quantitation. Once sequences are analyzed,
the results can be visualized to allow comparisons between samples and
different conditions summarized in the metadata.
**Some things to remember about the platform**
- You must be a registered CyVerse user to use Purple Line (register
for a CyVerse account at [user.cyverse.org](https://user.cyverse.org/){target=_blank}).
- The Purple line was designed to make microbiome/eDNA data analysis
"simple". However, we ask that users very carefully and
thoughtfully decide what "jobs" they want to submit.
- A single Purple Line project may take hours to process since HPC
computing is subject to queues which may support hundreds of other
jobs. These systems also undergo regular maintenance and are subject
to periodic disruption.
- DNA Subway implements the [QIIME 2](https://qiime2.org/){target=_blank} software. This software is in continual
development. Our version may not be the most current, and our
documentation and explanation is not meant to replace the full [QIIME 2 documentation](https://docs.qiime2.org/2018.8/){target=_blank}.
- We have made design decisions to create a straightforward
classroom-friendly workflow. While this Subway Line does not have
all possible features of QIIME 2, we purpose to cover important
concepts behind microbiome and eDNA analysis.
- You may work with up to 96 samples (e.g. 192 paired files or 96 single read files) in a Purple line project.
------------------------------------------------------------------------
!!! warning "Sample Data: How to use provided sample data"
In this guide, we will use a microbiome dataset (*"ubiome-test-data"*) collected from various
water sources in Montana (down-sampled and de-identified).Where
appropriate, a note (in this orange colored background) in the
instructions will indicate which options to select to make use of this
provided dataset.
### DNA Subway Purple Line - Metadata file and Sequencing Prerequisites
If you are generating data for a project (i.e. sequencing samples), you
will need to provide the sequencing data (fastq files) as well as a
metadata file that describes the data contained in these sequencing
files. This metadata must conform to strict guidelines, or analyses will
fail. QIIME 2 metadata is stored in a TSV (tab-separated values) file.
These files typically have a .tsv or .txt file extension, though it
doesn't matter to QIIME 2 what file extension is used. TSV files are
simple text files used to store tabular data, and the format is
supported by many types of software, such as editing, importing, and
exporting from spreadsheet programs and databases. Thus, it's usually
straightforward to manipulate QIIME 2 metadata using the software of
your choosing. If in doubt, we recommend using a spreadsheet program
such as Microsoft Excel or Google Sheets to edit and export your
metadata files.
**Handling Project Metadata**
Before you create your project, you will have generated metadata (as
described above) for your project. You have two options for preparing
this metadata to ensure that it conforms to the required QIIME2
parameters. The file must be validated (which you can do on your own or
using Subway). If there are errors in your file (this is common), they
must be fixed.
!!! tip "Formatting Your Metadata"
**Leading and trailing whitespace characters**
If any cell in the metadata contains leading or trailing whitespace
characters (e.g. spaces, tabs), those characters will be ignored when
the file is loaded. Thus, leading and trailing whitespace characters
are not significant, so cells containing the values 'gut' and ' gut' are equivalent. This rule is applied before any other rules
described below
**ID column**
The first column MUST be the ID column name (i.e. ID header) and the
first line of this column should be #SampleID or one of a few
alternative.
- Case-insensitive: id; sampleid; sample id; sample-id; featureid;
feature id; feature-id.
- Case-sensitive: #SampleID; #Sample ID; #OTUID; #OTU ID;
sample_name
**Sample IDs**
For the sample IDs, there are some simple rules to comply with QIIME 2
requirements:
- IDs may consist of any Unicode characters, with the exception
that IDs must not start with the pound sign (#), as those rows
would be interpreted as comments and ignored. IDs cannot be
empty (i.e. they must consist of at least one character).
- IDs must be unique (exact string matching is performed to detect
duplicates).
- At least one ID must be present in the file.
- IDs cannot use any of the reserved ID column names (the sample
ID names, above).
- The ID column can optionally be followed by additional columns
defining metadata associated with each sample or feature ID.
Metadata files are not required to have additional metadata
columns, so a file containing only an ID column is a valid QIIME
2 metadata file.
**Column names**
- May consist of any Unicode characters.
- Cannot be empty (i.e. column names must consist of at least one
character).
- Must be unique (exact string matching is performed to detect
duplicates).
- Column names cannot use any of the reserved ID column names.
**Column values**
- May consist of any Unicode characters.
- Empty cells represent missing data. Note that cells consisting
solely of whitespace characters are also interpreted as missing
data.
QIIME 2 currently supports categorical and numeric metadata columns.
By default, QIIME 2 will attempt to infer the type of each metadata
column: if the column consists only of numbers or missing data, the
column is inferred to be numeric. Otherwise, if the column contains
any non-numeric values, the column is inferred to be categorical.
Missing data (i.e. empty cells) are supported in categorical columns
as well as numeric columns. For more details, and for how to define
the nature of the data when needed, see the [QIIME 2 metadata documentation](https://docs.qiime2.org/2019.10/tutorials/metadata/){target=_blank}.
**Working with an existing metadata file**
!!! tip
If you have your own metadata file, it will still need to be validated
once uploaded to DNA Subway.
Using a spreadsheet editor, create a metadata sheet that provides
descriptions of the sequencing files used in your experiment. Export
this file as a tab-delimited **.txt** or **.tsv** file. following
the QIIME 2 metadata documentation](https://docs.qiime2.org/2019.10/tutorials/metadata/) recommendations. (Optional: if you using your own metadata file you can validate it using DNA Subway and or online QIIME2 validator [Keemei](https://keemei.qiime2.org/){target=_blank}).
!!! tip
See an example metadata file used for our sample data here: [metadata file](http://datacommons.cyverse.org/browse/iplant/home/shared/cyverse_training/platform_guides/dna_subway/purple_line/mappingfile.xlsx){target=_blank}. Click the **Download** button on the
linked page to download and examine the file. (**Note**: This is an
Excel version of the metadata file, you must save Excel files as .TSV
(tab-separated) to be compatible with the QIIME 2 workflow.)
**Creating a metadata file using DNA Subway**
See [DNA Subway Purple Line - Metadata and QC](#dna-subway-purple-line-metadata-and-qc) section C.
### DNA Subway Purple Line - Create a Microbiome Analysis Project
**A. Create a project in Subway**
1. Log-in to DNA Subway (unregistered users may NOT use Purple Line,
register for a CyVerse account at [user.cyverse.org](https://user.cyverse.org/){target=_blank}.
2. Click the purple square ("Microbiome Analysis") to begin a
project.
3. For 'Select Project Type' select either **Single End Reads** or
**Paired End Reads**
!!! warning "Sample Data"
*"ubiome-test-data"* dataset: Select **Single End Reads**
4. For 'Select File Format' select the format the corresponds to
your sequence metadata.
!!! warning "Sample Data"
*"ubiome-test-data"* dataset: Select **Illumina Casava 1.8**
!!! tip
Typically, microbiome/eDNA will be in the form of multiplexed FastQ sequences. We support the following formats:
- [Illumina Casava 1.8](http://illumina.bioinfo.ucr.edu/ht/documentation/data-analysis-docs/CASAVA-FASTQ.pdf/at_download/file){target=_blank}
5. Enter a project title, and description; click **Continue**.
**B. Upload read data to CyVerse Data Store**
The sequence read files used in these experiments are too large to
upload using the Subway interface. You must upload your files (**Note**: Only `.fastq.gz` files are accepted) directly to the CyVerse Data Store:
1. Upload your
- FASTQ sequence reads; **Note**: Only `.fastq.gz` files are accepted.
- Sample metadata file (.tsv or .txt formatted according to [QIIME 2 Metadata documentation](https://docs.qiime2.org/2019.10/tutorials/metadata/){target=_blank}) to
the CyVerse Data Store using Cyberduck. See instructions: [CyVerse Data Store Guide](https://unm-carc.github.io/cyverse/data-store/sftp/cyberduck/).
(Optional: You can edit and change metadata using the Subway
interface in the [Manage data]step once the project is
created.)
------------------------------------------------------------------------
### DNA Subway Purple Line - Metadata and QC
**A. Select files using Manage Data**
1. Click on the 'Manage data stop: this opens a window where you can
add your FASTQ (up to 192 paired files or 96 single read files) and metadata files. Click
**+Add from CyVerse** to add the
FASTQ files uploaded to the CyVerse Data Store. Select your files
and then click **Add selected files** or **Add all FASTQ files in this directory** as appropriate. **Note**: Only `.fastq.gz` files are accepted.
!!! warning "Sample Data"
*"ubiome-test-data"* dataset: Navigate to: Shared Data > SEPA_microbiome_2016 > **ubiome-test-data** and click **Add all FASTQ files in this directory**
2\. To add your metadata file you may use one of three options:
- *Add from CyVerse*: Add a metadata file you have uploaded to CyVerse Data store
- *Upload locally*: Directly upload a metadata file from your local computer
- *Create New*: Create a new metadata file using DNA Subway
!!! tip "Creating a metadata file using DNA Subway"
You can create a metadata file using DNA Subway. Creating the file
step-by-step will help you to avoid metadata errors. Be sure you
have consulted the [QIIME 2 documentation](https://docs.qiime2.org/2019.10/tutorials/metadata/){target=_blank} so you can anticipate what the required fields
are. To use this feature under in the 'Manage data' step under
'Metadata Files' click **Create new**
**Sample IDs and adding/removing samples**
These are unique IDs for each of your samples.
All metadata files must have a column called **#SampleID**. Click
**+Add samples** to add additional
rows. In the Subway form, these will be unique, arbitrary names
(roughly corresponding to well-positions on a 96-well microplate).
You can change these (including pasting in sample names from an
existing spreadsheet).
{width="450px" height="250px"}
Right-clicking on a row number allows you to remove or insert rows.
{width="450px" height="250px"}
**Adding columns, managing sample descriptions and data types**
The very **last** column must be a sample description. You can click
the arrow on the right of this column to add a new column (which
will be added to the left). Column names must be unique, must not be
empty, cannot contain whitespace, can contain a maximum of 32
characters, cannot match a reserved column name. Notice that when
you click on a column name it is colored -pink for columns that have
numeric data (e.g. measurements) and cyan for everything else (e.g.
categorical descriptions in the form of words (i.e. strings)).
Clicking a column name will allow you to change its type.
{width="450px" height="250px"}
**Handling errors**
If you violate one of the rules for metadata formatting, the entry
will turn red. Consult the help and or the [QIIME 2 documentation](https://docs.qiime2.org/2019.10/tutorials/metadata/){target=_blank} to correct the error.
{width="450px" height="250px"}
Click **Save** to save your
metadata file, and close the window.
!!! warning "Sample Data"
*"ubiome-test-data"* dataset:
Click **Add from CyVerse**Navigate to: Shared Data > SEPA_microbiome_2016 > **ubiome-test-data**
Select the **mappingfile_MT_corrected.tsv** and then click **Add selected files**.
3\. As needed, you can edit or rename your metadata file. Before
proceeding, you must validate your metadata file. To validate,
click the "validate" link to the right of the metadata file you
wish to check. Once the validation completed, click `Run` to proceed. If you have
errors, you will be presented with an `Edit` button so that you can return to the file and
edit.
**B. Demultiplex reads**
At this step, reads will be grouped according to the sample metadata.
This includes separating reads according to their index sequences if
this was not done prior to running the Purple Line. For demultiplexing
based on index sequences, the index sequences must be defined in the
metadata file.
!!! note
Even if your files were previously demultimplexed (as will generally
be the case with Illumina data) you must still complete this step to
have your sequence read files appropriately associated with metadata.
1\. Click the 'Demultiplex reads' and choose a number of reads to sample. When the job has completed click *Demultiplexing Summary* to view your results. In 'Random sequences to sample for QC', enter a value (1000 is recommended),
!!! warning "Sample Data"
*"ubiome-test-data"* dataset: Use the default of 1000 sequences
2\. When demultiplexing is complete, you will generate a file (.qzv) click this link to view a visualization and statistics on the sequence and metadata for this project.
!!! tip
Several jobs on Purple Line will take several minutes to an hour
to complete. Each time you launch one of these steps you will get
a Job ID. You can click the **View job info** button to see a detailed status and
diagnostic/error messages. If needed There is a *stop this job* link at the bottom of the info page to cancel a
job.
!!! note
**QIIME2 Visualizations**
One of the features of QIIME 2 are the variety of visualizations
provided at several analysis steps. Although this guide will not
cover every feature of every visualization, here are some
important points to note.
**QIIME2 View**: DNA Subway uses the QIIME 2 View plugin to
display visualizations. Like the standalone QIIME 2
software, you can navigate menus, and interact with several
visualizations. Importantly, many files and visualizations
can be directly download for your use outside of DNA Subway,
including in report generation, or in your custom QIIME 2
analyses. You can view downloaded .qza or .qzv files at [view.qiime2.org](https://view.qiime2.org/){target=_blank}.
!!! tip "Quality Graphs Explained"
After demultiplexing, you will be presented with a visualization
that displays the following tables and graphs:
**Overview Tab**
- *Demultiplexed sequence counts summary*: For each of the
fastq files (each of which may generally correspond to a
single sample), you are presented with comparative
statistics on the number of sequences present. This is
followed by a histogram that plots number of sequences by
the number of samples.
- *Per-sample sequence counts*: These are the actual counts
of sequences per sample as indicated by the sample names
you provided in your metadata sheet.
{width="450px" height="250px"}
**Interactive Quality Plot**
This is an interactive plot that gives you an average quality
(y-axis) by the position along the read (x-axis). This box plot
is derived from a random sampling of a subset of sequences. The
number of sequences sampled will be indicated in the plot
caption. You can use your mouse drag and zoom in to regions on
the plot. Double-click your mouse to zoom out.
{width="450px" height="250px"}
3\. Click the "Interactive Quality Plot" tab to view a histogram of sequence quality. Use this plot at the tip below to determine a location to trim.
!!! tip "Tips on trimming for sequence quality"
On the Interactive Quality Plot you are shown an histogram, plotting
the average quality (x axis) [Phred score](https://en.wikipedia.org/wiki/Phred_quality_score){target=_blank} vs. the position on the read (y axis) in base pairs for a **subsample** of reads.
**Zooming to determine 3' trim location**
Click and drag your mouse around a collection of base pair positions
you wish to examine. Clicking on a given histogram bar will also
generate a text report and metrics in the table below the chart. Using
these metrics, you can choose a position to trim on the right side
(e.g. 3' end of the sequence read). The 5' (left trim) is specific
to your choice of primers and sequencing adaptors (e.g. the sum of the
adaptor sequence you expect to be attached to the 5' end of the
read). Poor quality metrics will generate a table colored in red, and
those base positions will also be colored red in the histogram.
Double-clicking will return the histogram to its original level of
zoom.
**Example plots**
It is important to maximize the length of the reads while minimizing
the use of low quality base calls. To this end, a good guideline is to
trim the right end of reads to a length where the 25th percentile is
at a quality score of 25 or more. However, the length of trimming will
depend on the quality of the sequence, so you may have to use a lower
quality threshold to retain enough sequence for informative sequence
searches and alignments. This may require multiple runs of the
analysis to find the optimal trim length for your data.
*Quality drops significantly at base 35*
{width="400px" height="250px"}
*Improved quality sequence*
{width="400px" height="250px"}
**C. Use DADA2 for Trimming and Error-correction of Reads**
It is important to only work with high quality data. This step will
generate a sequence quality histogram which can be used to determine
parameter for trimming.
1. Click 'DADA2' and choose the metadata file corresponding to the
samples you wish to analyze. Then choose values for trimming of
the reads. For "trimLeft" (the position starting from the left
you wish to trim) and "TruncLen" (this is the position where
reads should be trimmed, truncating the 3' end of the read. Reads
shorter than this length will be discarded). Finally, click
**Trim reads**.
!!! warning "Sample Data"
*"ubiome-test-data"* dataset: Based on the histogram for our sample, we recommend the following parameters:
- **trimLeft: 17** (this is specific to primers and adaptors in
this experiment)
- **TruncLen: 200** (this is where low quality sequence begins, in
this case because our sequence length is lower than the expected
read length)
**D. Check Results of Trimming** Once trimming is complete, the
following outputs are expected:
1. Click on DADA2 and then click on the links in the *Results* table
to examine results.
- **Trim Table** (*Metric summary*, *Frequency per sample*,
*Frequency per feature*): Summarizes the dataset post-trimming
including the number of samples and the number of features per
sample. The "Interactive Sample Detail" tab contains a sampling
depth tool that will be used in computation of the core matrix.
!!! note
**You will use the maximum frequency value for the Alpha
rarefaction step** So you may wish to record this value now for
the DNA Subway 'Clustering sequences' step.
- **Stats**: Sequencing statistics for each of the sample IDs
described in the original metadata file.
- **Representative Sequences** (*Sequence Length Statistics*,
*Seven-Number Summary of Sequence Lengths*, *Sequence Table*):
This table contains a listing of features observed in the sequence
data, as well as the DNA sequence that defines a feature. Clicking
on the DNA sequence will submit that sequence for BLAST at NCBI in
a separate browser tab.
The feature table contains two columns output by DADA2. DADA2
(Divisive Amplicon Denoising Algorithm 2) determines what sequences
are in the samples. DADA2 filters the sequences and identifies
probable amplification or sequencing errors, filters out chimeric
reads, and can pair forward and reverse reads to create the best
representation of the sequences actually found in the samples and
eliminating erroneous sequences.
- **Feature ID**: A unique identifier for sequences.
- **Sequence**: A DNA Sequence associated with each identifier.
Clicking on any given sequence will initiate at BLAST search on the
NCBI website. Click "View report" on the BLAST search that opens in
a new web browser tab to obtain your results. Keep in mind that if
your sequences are short (due to read length or trimming) many BLAST
searches may not return significant results.
!!! tip
Although the term "feature" can (unfortunately) [have many meanings](https://forum.qiime2.org/t/what-is-a-feature-exactly/2201){target=_blank} as used by the
QIIME2 documentation, unless otherwise noted in this documentation
it can be thought of as an OTU ([Operational Taxonomic Unit](https://en.wikipedia.org/wiki/Operational_taxonomic_unit){target=_blank}); another substitution for the word
species. OTU is a convenient and common terminology for referring to
an unclassified or undetermined species. Ultimately, we are
attempting to identify an organism from a sample of DNA which may
not be informative enough to reach a definitive conclusion.
!!! tip
If you want to redo the DADA2 step with different parameters, click
the "New Job" tab on the upper left of a DADA2 window to submit a
new job. New jobs appear as tabs on Subway steps that are typically
run several times. You can go back an see these jobs which are
labeled with a job number.
{width="450px" height="250px"}
------------------------------------------------------------------------
### DNA Subway Purple Line - Alpha Rarefaction/Clustering Sequences
**A. Alpha rarefaction**
At this step, you can visualize summaries of the data. A feature table
will generate summary statistics, including how many sequences are
associated with each sample. **Note that sample depth is limited to 100,000.**
1. Click on 'Alpha rarefaction'. Select "run" and designate the
minimum and maximum rarefaction depth. A minimum value should be
set at 1. The maximum value is specific to your data set. The
maximum value is specific to your data set. To determine what the
maximum value should be set to, open the "Trim Table" from the
"DADA2" step. You may not choose a value that is greater than
the maximum frequency per sample. In general, choosing a value
that is somewhere around the median frequency seems to work well,
but you may want to increase that value if the lines in the
resulting rarefaction plot don't appear to be leveling out, or
decrease that value if you seem to be losing many of your samples
due to low total frequencies closer to the minimum sampling depth
than the maximum sampling depth. Identify the maximum Sequence
Count value and enter that number as the maximum value. Click
**Submit Job**.
!!! note
Since you may want to try Alpha rarefaction using different
combinations of results from DADA2 trimming and your choice of
rarefaction depths, your trim (DADA2) jobs are displayed on the
left, and each new Alpha rarefaction setting will appear as a tab on
the top.
{width="450px" height="250px"}
!!! warning "Sample Data"
*"ubiome-test-data"* dataset:
We recommend the following parameters:
- **Min. rarefaction depth**: 1
- **Max. rarefaction depth**: 2938
2. Under 'Results' click on **Alpha Rarefaction Plot** to view the results.
??? tip "Navigating Alpha Rarefaction graphs"
**Alpha rarefaction** generates an interactive plot of species
diversity by sampling depth by the categorical samplings described
in your sample metadata. You can use dropdown menus to change
metrics/conditions displayed and also export data as a CSV file.
{width="450px" height="250px"}
------------------------------------------------------------------------
### DNA Subway Purple Line - Calculate Core Metrics/Alpha and Beta Diversity
At this stop, you will examine *Alpha Diversity* (the diversity of
species/taxa present within a single sample) and *Beta Diversity* (a
comparison of species/taxa diversity between two or more samples).
- Alpha diversity answers the question - "How many species are in a
sample?"
- Beta diversity answer the question - "What are the differences in
species between samples?"
**A. Calculate core metrics**
1. Click on 'Core metrics' and then click the "run" link. Choose
a sampling depth based upon the "Sampling depth" tool (described
in Section D Step 1, in the *Trim Table* output; *Interactive
Sample Detail* tab). Choose an appropriate classifier (see
comments in the tip below) and click **Submit job**.
!!! tip "Choosing Core metrics parameters"
*Sampling Depth*
In downstream steps, you will need to choose a sampling depth for
your sample comparisons. You can choose by examining the table
generated at the **Trim reads** step. In the *Trim Table* output,
*Interactive Sample Detail* tab, use the "Sampling depth" tool to
explore how many sequences can be sampled during the Core matrix
computation. As you slide the bar to the right, more sequences are
sampled, but samples that do not have this many sequences will be
removed during analysis. The sampling depth affects the number of
sequences that will be analyzed for taxonomy in later steps: as the
sampling depth increases, a greater representation of the sequences
will be analyzed. However, high sampling depth could exclude
important samples, so a balance between depth and retaining samples
in the analysis must be found.
*Classifier*
Choose a classifier pertaining to your experiment type.
- **Microbiome** choose **Greengenes (515F/806R)** or **Greengenes
(full sequences)** or **Sliva (16S rRNA)** classifier
- **eDNA** experiment with marine fishes you may elect to choose
the **Fish 12S/ecoPrimer** classifier
!!! warning "Sample Data"
*"ubiome-test-data"* dataset:
We recommend the following parameters:
- **Sampling Depth**: 3000
- **Classifier**: Grenegenes (full sequences)
**B. Examine alpha and beta diversity**
2. When core metrics is complete, you should generate several
visualization results. Click each of the following to get access:
- **Alpha Diversity:**
- ***Pielou's Evenness***
- Alpha Correlation: Measure of community evenness using correlation tests
- Group Significance: Analysis of differences between features across group
- ***Faith's Phylogenetic Diversity***
- Alpha Correlation: Faith Phylogenetic Diversity (a measure of community richness) with correlation tests
- Group Significance: Faith Phylogenetic Diversity ( a measure of community richness)
- **Beta Diversity:**
- ***Bray-Curtis Distance***
[Bray-Curtis](https://en.wikipedia.org/wiki/Bray%E2%80%93Curtis_dissimilarity){target=_blank} is a metric for describing the dissimilarity of species in an ecological sampling.
- Bioenv: Bray-Curtis test metrics
- Emperor: Interactive PCoA plot of Bray-Curtis metrics
- ***Jaccard Distance***
- Emperor: Interactive PCoA plot calculated by [Jaccard](https://en.wikipedia.org/wiki/Jaccard_index){target=_blank} similarity index.
- ***Unweighted UniFrac Distance***
[UniFrac](https://en.wikipedia.org/wiki/UniFrac){target=_blank} is a metric for describing the similarity of a biological community, taking into account the relatedness of community members.
- Bioenv: UniFrac test metrics
- Emperor: Unweighted interactive PCoA plot
- ***Weighted UniFrac Distance***
Unweighted UniFrac removes the effect of low-abundance features in the calculation of principal components.
- Emperor: Weighted interactive PCoA plot of UniFrac.
!!! tip
**Emperor Plots**
These plots are all interactive three-dimensional plots of an
analysis using [principal components](https://en.wikipedia.org/wiki/Principal_component_analysis){target=_blank}.
**Customization**
You can customize Emperor plots, including altering the color of
and shape points, axes, and other parameters. You can also export
images from this visualization.
{width="550px" height="300px"}
**Bioenv**
These plots are tables of tests and descriptive metrics.
**C. Taxonomic Diversity:**
Taxonomic diversity is at the heart of many analyses. We suggest
consulting the QIIME [taxonomy overview](https://docs.qiime2.org/2019.1/tutorials/overview/#taxonomy){target=_blank} for a detailed explanation of how QIIME2 calculates taxonomy and additional features of QIIME2 you may wish to
explore beyond the functionalities DNA Subway has included.
- *Bar Plots*
- An interactive stacked bar plot of species diversity.
Dropdown menus allow you to color by seven taxonomic
levels 1) kingdom, 2) phylum,
3) class, 4) order, 5), family, 6) genus, 7) species. Plots
can be further arranged/filtered/sorted accoridng to
characteristics in the sample metadata. You may also
download images and data used to create the barpot
visualization.
{width="550px" height="300px"}
- *Taxonomy*
- A table indicating the identified "features", their taxa,
and an indication of confidence.
You can download and interact with any of the available plots.
**D. Calculate differential abundance**
1. Click on the 'Differential abundance' stop. Then click on the
"Submit new "Differential abundance" job" link. Choose a
metadata category to group by, and a level of taxonomy to
summarize by. Then click **submit job**.
!!! warning "Sample Data"
*"ubiome-test-data"* dataset:
We recommend the following parameters:
- **Group data by**: CollectionMethod
- **Level of taxonomy to summarize**: 5
!!! tip
Download the provided CSV files so that you can generate customized plots.
------------------------------------------------------------------------
### DNA Subway Purple Line - Visualize data with PiCrust and PhyloSeq
**Under Development**
------------------------------------------------------------------------
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/edu/tutorials/dna_subway_guide.md){target=_blank} (last source update 2025-03-14), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.
---8<--- https://unm-carc.github.io/cyverse/developers/manuals/
---
title: "Developer manuals"
description: "Where to find documentation for each CyVerse platform and API, from the Data Store and Discovery Environment to Terrain, Tapis, and CACAO."
type: Reference
tags:
- Developers
- API
- Documentation
generated:
by: "claude/opus-5"
at: "2026-09-11T00:00:00Z"
sources:
- id: cyverse-learning-materials
resource: "https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/dev/manuals.md"
title: "CyVerse Learning Materials: docs/dev/manuals.md"
author: "team:cyverse"
last_modified: "2025-03-16T09:25:29-07:00"
---
# Developer manuals
Table: Documentation websites for each CyVerse platform
| Platform | Purpose | Link |
|----------|---------|------|
| Data Store | data hosting and storage | [Data Store](https://unm-carc.github.io/cyverse/data-store/overview/) |
| Discovery Environment UI | analysis and workbench | [Discovery Environment](https://unm-carc.github.io/cyverse/discovery-environment/overview/) |
| Discovery Environment API | API access | [DE API Manual](https://docs.cyverse.org/api/terrain/){target=_blank} |
| CACAO Jetstream2 | Cloud Services | [CACAO Documentation](https://docs.jetstream-cloud.org/ui/cacao/overview/){target=_blank} |
| Atmosphere (Deprecated) | Cloud Services | [Atmosphere Manual](https://cyverse.atlassian.net/wiki/spaces/atmman/overview){target=_blank} |
| BISQUE (Deprecated) | Image Analysis | [BiSQUE Guide](https://cyverse.atlassian.net/wiki/spaces/BIS/overview){target=_blank} |
| DNA Subway (Deprecated) | Genomics Education | [DNA Subway Guide](https://unm-carc.github.io/cyverse/education/tutorials/dna-subway/) |
| TAPIS | TACC APIs | [TAPIS](https://tapis-project.org/){target=_blank} |
!!! info "Under the hood"
For the Discovery Environment's REST API and for contributing to CyVerse's own services, start with [Endpoint index](https://docs.cyverse.org/api/endpoint-index/){target=_blank} and [Developer guide](https://docs.cyverse.org/development/developer-guide/){target=_blank} in the CyVerse Developer Documentation.
Adapted from [CyVerse Learning Materials](https://github.com/CyVerse-learning-materials/learning-materials-home/blob/b7392d21be4fd6a67d051847f98c81a18e058a12/docs/dev/manuals.md){target=_blank} (last source update 2025-03-16), CC BY 4.0. Spotted a problem? [Open an issue](https://github.com/UNM-CARC/cyverse/issues){target=_blank}.