# How to iterate over groupTuple data for processes?

**URL:** <https://community.seqera.io/t/how-to-iterate-over-grouptuple-data-for-processes/396>\
**Category:** Ask for help\
**Created:** [January 16, 2024, 12:42am UTC](https://community.seqera.io/t/how-to-iterate-over-grouptuple-data-for-processes/396 "2024-01-16T00:42:12Z")\
**Posts on this page:** 9\
**Page:** 1

<div class="post-metadata">

**Author:** ![complexgenome](https://dub1.discourse-cdn.com/flex013/user_avatar/community.seqera.io/complexgenome/32/392_2.png) [@complexgenome](https://community.seqera.io/u/complexgenome)\
**Post date:** [January 16, 2024, 12:42am UTC](https://community.seqera.io/t/how-to-iterate-over-grouptuple-data-for-processes/396/1 "2024-01-16T00:42:12Z")

</div>

Hello there,

I’ve a data as below (without header):

| Patient\_ID | Sample name | DNA\_N\_R1 | DNA\_N\_R2 | DNA\_T\_R1 | DNA\_T\_R2 | RNA\_T\_R1 | RNA\_T\_R2 |
| --- | --- | --- | --- | --- | --- | --- | --- |
| patient1 | patient1-3 | DNA1\_N\_R1 | DNA1\_N\_R2 | DNA\_T\_T01\_R1 | DNA\_T\_T01\_R2 | RNA\_T\_T01\_R1 | RNA\_T\_T01\_R2 |
| patient1 | patient1-4 | DNA1\_N\_R1 | DNA1\_N\_R2 | DNA\_T\_T02\_R1 | DNA\_T\_T02\_R2 | RNA\_T\_T02\_R1 | RNA\_T\_T02\_R2 |
| patient1 | patient1-5 | DNA1\_N\_R1 | DNA1\_N\_R2 | DNA\_T\_T03\_R1 | DNA\_T\_T03\_R2 | RNA\_T\_T03\_R1 | RNA\_T\_T03\_R3 |
| patient2 | patient2-5 | DNA2\_N\_R1 | DNA2\_N\_R2 | DNA2\_T\_T03\_R1 | DNA2\_T\_T03\_R2 | RNA2\_T\_T03\_R1 | RNA2\_T\_T03\_R3 |

Each of the columns from 3rd are a DNA/RNA compressed raw FASTQ (sequenced) files. For brevity I’ve not put .fastq.gz, however, the idea is to have paired files for eventual data processing.

Goal: I’d like to run per-patient, per row processes.

Steps/code:

```auto
    Channel.fromPath(file("input_timestamp.csv"))
    .splitCsv(sep: ',')
    .groupTuple().map { row -> 
            // Extract relevant information
            def patient_info = row[0]
            def sample_info=row[1]
            def normal_reads = tuple((row[2]),(row[3]))
            def tumor_reads = tuple((row[4]), (row[5]))
            def rna_reads = tuple((row[6]), (row[7]))        

            // Return a map with the processed information
            return [patient: patient_info, sample:sample_info,normal: normal_reads, tumor: tumor_reads, rna: rna_reads]
        }
        .set { samples_grouped }

```

1. 

Now, how do I do what? How do I iterate over each row?

1. When I print these I think I doubt the tuple/pair/print. For e.g. when I print the above `samples_grouped` as `samples_grouped.view { "$it" }` I get output as:

> [patient:patient1, sample:[sample1-1, sample1-2], normal:[[DNA\_N\_R1, DNA\_N\_R1], [DNA\_N\_R2, DNA\_N\_R2]], tumor:[[DNA\_T\_T01\_R1, DNA\_T\_T02\_R1], [DNA\_T\_R2, DNA\_T\_T02\_R2]], rna:[[RNA\_T\_01\_R1, RNA\_T\_02\_R1], [RNA\_T\_01\_R2, RNA\_T\_02\_R2]]]
> 
> [patient:patient2, sample:[sample2], normal:[[DNA\_N\_R1], [DNA1\_N\_R2]], tumor:[[DNA\_T\_T03\_R1], [DNA\_T\_T03\_R2]], rna:[[RNA1\_T\_01\_R1], [RNA\_T\_01\_R2]]]

I mean, if I look at row 1 for patient1, the normal is [DNA\_N\_R1, DNA\_N\_R1]  
Is this correct? How do I access normal DNA\_N\_R1 with DNA\_N\_R2?

Thank you in advance.

---

<div class="post-metadata">

**Author:** ![mribeirodantas](https://dub1.discourse-cdn.com/flex013/user_avatar/community.seqera.io/mribeirodantas/32/235_2.png) [@mribeirodantas](https://community.seqera.io/u/mribeirodantas)\
**Post date:** [January 16, 2024, 1:36pm UTC](https://community.seqera.io/t/how-to-iterate-over-grouptuple-data-for-processes/396/2 "2024-01-16T13:36:15Z")

</div>

The short answer is that you don’t need the `groupTuple`. You can keep the patient information in the channel element so that you’ll always know what patient is a sample from. The only difference from my input CSV file compared to yours is that I added the path to the file (`/foo/bar/DNA_N_R2.fast.gz` instead of `DNA_N_R2`). My Nextflow script file (without `groupTuple`):

```Groovy
process PROCESS_SAMPLES {
  debug true

  input:
  tuple val(patient_id),
        val(sample_id),
        path(normal_reads),
        path(tumor_reads),
        path(rna_reads)

  output:
  stdout

  script:
  """
  echo ---------------------------------------------------
  echo Doing something on ${sample_id} from ${patient_id}
  echo Normal reads: 1- ${normal_reads[0]} 2- ${normal_reads[1]}
  echo Tumor reads: 1- ${tumor_reads[0]} 2- ${tumor_reads[1]}
  echo RNA reads: 1- ${rna_reads[0]} 2- ${rna_reads[1]}
  """
}

workflow {
  Channel
    .fromPath(file("input.csv"))
    .splitCsv(sep: ',')
    .map { row ->
      // Extract relevant information
      def patient_info = row[0]
      def sample_info=row[1]
      def normal_reads = tuple((row[2]),(row[3]))
      def tumor_reads = tuple((row[4]), (row[5]))
      def rna_reads = tuple((row[6]), (row[7]))

      // Return a map with the processed information
      return [patient: patient_info, sample:sample_info, normal: normal_reads, tumor: tumor_reads, rna: rna_reads]
    }
    .set { samples }
  PROCESS_SAMPLES(samples)
}

```

The output:

 ![Captura de Tela 2024-01-16 às 10.35.12](https://europe1.discourse-cdn.com/flex013/uploads/seqera/original/1X/b864eb5355ed85ac55f188cff06efd95a40a7d6b.jpeg)

Two sample working directories, to show you the files are being correctly staged:

 ![Captura de Tela 2024-01-16 às 10.35.50](https://europe1.discourse-cdn.com/flex013/uploads/seqera/original/1X/7d8165b709f6c6c1b2f8adf66332b77f258e5638.jpeg)

---

<div class="post-metadata">

**Author:** ![complexgenome](https://dub1.discourse-cdn.com/flex013/user_avatar/community.seqera.io/complexgenome/32/392_2.png) [@complexgenome](https://community.seqera.io/u/complexgenome)\
**Post date:** [January 16, 2024, 3:57pm UTC](https://community.seqera.io/t/how-to-iterate-over-grouptuple-data-for-processes/396/3 "2024-01-16T15:57:11Z")

</div>

@mribeirodantas

Thank you for your reply.

How do I have nextflow accept path even if it doesn’t exist? I get error as:

> Caused by:  
> Not a valid path value type: nextflow.util.ArrayBag ([/data1/daphni2/clinicalfq/WES/sema4/clinical/CLN-22087270-DNA-N-0\_HJFJTDSX3\_S22\_L002\_R2.fastq.gz]

or,

> Caused by:  
> Not a valid path value type: nextflow.util.ArrayBag ([NA])

There will be times when patient’s either RNA or WES will only be available.

For e.g., there will be analysis where there is no DNA (normal and tumor) but only RNA, so it will be NA. I get error when file are not present using path.

My code is structured as:  
1- user provides a CSV file  
2- user select which analysis to run: WES or RNA. user can select both as well, in that case I run DNA and RNA workflows both.  
based on selection, I used to pass samples.rna or samples.normal, samples.tumor.

How do I achieve with your code/snippet?

How do I make nextflow accept path for files that is not present?  
Or, how do I only pass: patient, sample, normal-reads, tumor-reads and/or rna-reads based on user choice.

---

<div class="post-metadata">

**Author:** ![mribeirodantas](https://dub1.discourse-cdn.com/flex013/user_avatar/community.seqera.io/mribeirodantas/32/235_2.png) [@mribeirodantas](https://community.seqera.io/u/mribeirodantas)\
**Post date:** [January 16, 2024, 4:48pm UTC](https://community.seqera.io/t/how-to-iterate-over-grouptuple-data-for-processes/396/4 "2024-01-16T16:48:14Z")

</div>

> [@complexgenome](#):
>
> How do I have nextflow accept path even if it doesn’t exist? I get error as:

The path did not exist in my example. I have no `/foo/bar` in my machine 😅 (and I got no errors)

> [@complexgenome](#):
>
> Caused by:  
> Not a valid path value type: nextflow.util.ArrayBag ([/data1/daphni2/clinicalfq/WES/sema4/clinical/CLN-22087270-DNA-N-0\_HJFJTDSX3\_S22\_L002\_R2.fastq.gz]

Can you share a minimal reproducible example so that I can try to reproduce your errors here?

Also, if there are specific situations (such as NA) in your problem, you have to handle them. Feel free to open a new topic if you think it strays too much from this one.

---

<div class="post-metadata">

**Author:** ![complexgenome](https://dub1.discourse-cdn.com/flex013/user_avatar/community.seqera.io/complexgenome/32/392_2.png) [@complexgenome](https://community.seqera.io/u/complexgenome)\
**Post date:** [January 16, 2024, 6:24pm UTC](https://community.seqera.io/t/how-to-iterate-over-grouptuple-data-for-processes/396/5 "2024-01-16T18:24:10Z")

</div>

@mribeirodantas

I’m using: nextflow version 23.10.0

```auto

process DNA {

    input: tuple val(patient_id), val(sample_id), path(normal_reads), path(tumor_reads), path(rna_reads)
    output: stdout

    script:
  """
  echo ---------------------------------------------------
  echo Doing something on ${sample_id} from ${patient_id}
  echo Normal reads: 1- ${normal_reads[0]} 2- ${normal_reads[1]}
  echo Tumor reads: 1- ${tumor_reads[0]} 2- ${tumor_reads[1]}
  echo RNA reads: 1- ${rna_reads[0]} 2- ${rna_reads[1]}
  """
}

workflow {

Channel.fromPath(file("input_timestamp.csv"))
        .splitCsv(sep: ',')
        .map { row -> 
            // Extract relevant information
            def patient_info = row[0]
            def sample_info=row[1]
            def normal_reads = tuple((row[2]),(row[3]))
            def tumor_reads = tuple((row[4]), (row[5]))
            def rna_reads = tuple((row[6]), (row[7]))
            
            // Return a map with the processed information
            return [patient: patient_info, sample:sample_info,normal: normal_reads, tumor: tumor_reads, rna: rna_reads]
        }
        .set { samples }

DNA(samples).view()
}

```

Attached is error screenshot.

I used following data in a CSV file:

> patient1,sample1-1,DNA\_N\_R1,DNA\_N\_R2,DNA\_T\_T01\_R1,DNA\_T\_R2,RNA\_T\_01\_R1,RNA\_T\_01\_R2  
> patient1,sample1-2,DNA\_N\_R1,DNA\_N\_R2,DNA\_T\_T02\_R1,DNA\_T\_T02\_R2,RNA\_T\_02\_R1,RNA\_T\_02\_R2  
> patient2,sample2,DNA1\_N\_R1,DNA1\_N\_R2,DNA\_T\_T03\_R1,DNA\_T\_T03\_R2,RNA1\_T\_01\_R1,RNA\_T\_01\_R2

 ![Screenshot 2024-01-16 at 1.22.13 PM](https://europe1.discourse-cdn.com/flex013/uploads/seqera/original/1X/fd4b4b2e9cd925569980efed631c054e84b58b23.jpeg)

---

<div class="post-metadata">

**Author:** ![mribeirodantas](https://dub1.discourse-cdn.com/flex013/user_avatar/community.seqera.io/mribeirodantas/32/235_2.png) [@mribeirodantas](https://community.seqera.io/u/mribeirodantas)\
**Post date:** [January 17, 2024, 2:14pm UTC](https://community.seqera.io/t/how-to-iterate-over-grouptuple-data-for-processes/396/6 "2024-01-17T14:14:47Z")

</div>

> [@mribeirodantas](#):
>
> The only difference from my input CSV file compared to yours is that I added the path to the file (`/foo/bar/DNA_N_R2.fast.gz` instead of `DNA_N_R2`).

You must provide a path. You’re telling Nextflow this is a path, but providing simply a filename string.

---

<div class="post-metadata">

**Author:** ![complexgenome](https://dub1.discourse-cdn.com/flex013/user_avatar/community.seqera.io/complexgenome/32/392_2.png) [@complexgenome](https://community.seqera.io/u/complexgenome)\
**Post date:** [January 17, 2024, 3:54pm UTC](https://community.seqera.io/t/how-to-iterate-over-grouptuple-data-for-processes/396/7 "2024-01-17T15:54:57Z")

</div>

@mribeirodantas  
I see. Thank you for identifying that.

My follow up questions - 1) how to decide to when use multiMap v/s map?  
2) instead of unpacking the map in a process, would you suggest to use some for, each sort of thing to iterate over map?

Thank you again.

---

<div class="post-metadata">

**Author:** ![mribeirodantas](https://dub1.discourse-cdn.com/flex013/user_avatar/community.seqera.io/mribeirodantas/32/235_2.png) [@mribeirodantas](https://community.seqera.io/u/mribeirodantas)\
**Post date:** [January 17, 2024, 7:50pm UTC](https://community.seqera.io/t/how-to-iterate-over-grouptuple-data-for-processes/396/8 "2024-01-17T19:50:24Z")

</div>

> [@complexgenome](#):
>
> 1. how to decide to when use multiMap v/s map?

Let’s say you have a channel, want to apply a function to every element of this channel, and have a single channel as output/result. That’s a very common scenario and a common use case for the `map` channel operator.

```Groovy
Channel
  .of(1, 2, 3, 4)
  .map { it -> it*it }
  .view()

```

Output:

 ![CleanShot 2024-01-17 at 16.44.11@2x](https://europe1.discourse-cdn.com/flex013/uploads/seqera/original/1X/7bc8f01cc04dac279ca92db9ec876da1ecf7278e.png)

`multiMap`, on the other hand, is a channel operator that allows you to output a channel with channels inside. See the example below.

```Groovy
Channel
  .of(1, 2, 3, 4)
  .multiMap {
    squared: it * it
    one_more: it + 1
  }
  .one_more
  .view()

```

 ![CleanShot 2024-01-17 at 16.48.46@2x](https://europe1.discourse-cdn.com/flex013/uploads/seqera/original/1X/c4ceb6c8ab5442d158e81e0f329c6b87a64ce68c.png)

As you can see close to the end of the script, I picked `.one_more` to view, but I could have picked `.squared`. It’s important to be aware of this structure, as if you provide a channel of channels to a process, you’ll run into trouble.

As a reference, you can find more info about channel operators [here](https://www.nextflow.io/docs/latest/operator.html#). The foundational training material also covers channel operators [here](https://training.nextflow.io/basic_training/operators/).

> [@complexgenome](#):
>
> 1. instead of unpacking the map in a process, would you suggest to use some for, each sort of thing to iterate over map?

In Nextflow, it’s more natural to follow a functional paradigm using channel operators such as `map`. I have seen very few useful cases of `for`, for example. It’s usually not what you’re looking for.

---

<div class="post-metadata">

**Author:** ![system](https://dub1.discourse-cdn.com/flex013/user_avatar/community.seqera.io/system/32/2402_2.png) [@system](https://community.seqera.io/u/system)\
**Post date:** [January 24, 2024, 7:50pm UTC](https://community.seqera.io/t/how-to-iterate-over-grouptuple-data-for-processes/396/9 "2024-01-24T19:50:46Z")

</div>

This topic was automatically closed 7 days after the last reply. New replies are no longer allowed.
