﻿WEBVTT

00:00:07.716 --> 00:00:08.341
My name is

00:00:08.341 --> 00:00:11.594
doctor LNG and
I'm the Director of Informatics at r u p.

00:00:12.095 --> 00:00:15.473
When I tell others I'm a BI and
I'm often asked what is my informatics?

00:00:15.640 --> 00:00:16.599
My first response?

00:00:16.599 --> 00:00:20.979
It's complicated, but more specifically,
mathematics is an interdisciplinary field

00:00:20.979 --> 00:00:24.607
that pulls together biology, mathematics,
statistics and computer science

00:00:24.899 --> 00:00:27.986
to create computational pipelines
for dissecting biological data.

00:00:28.570 --> 00:00:30.071
Now, if that doesn't sound complicated,

00:00:30.071 --> 00:00:32.323
let me tell you
more about its applications.

00:00:32.323 --> 00:00:34.117
The variety of uses for bioinformatics.

00:00:34.117 --> 00:00:38.288
However, my work specifically focuses on
its key role in clinical genomic testing.

00:00:38.705 --> 00:00:41.833
Unlike other testing methodologies,
a high degree of computational support

00:00:41.833 --> 00:00:45.045
is required to interpret short
read next generation sequencing data

00:00:45.295 --> 00:00:49.007
to generate clinically actionable
reports, said another way.

00:00:49.049 --> 00:00:52.052
But for my tools at a supporting framework
that does the behind the scenes

00:00:52.052 --> 00:00:55.055
work to generate consistent
and reliable genetic test results.

00:00:55.305 --> 00:00:58.141
And then just my friend
pipeline is designed to convert raw

00:00:58.141 --> 00:01:02.479
sequencing data into medical content
by detecting genomic variance, calculating

00:01:02.479 --> 00:01:06.107
QC metrics, and providing annotations
that are interpreted by clinical staff.

00:01:09.736 --> 00:01:12.155
It takes a team of very from opticians
to put all this together

00:01:12.155 --> 00:01:15.158
by using a variety of off the shelf tools,
or even by generating novel

00:01:15.158 --> 00:01:18.161
algorithms to bring clinical utility
to sequence analysis.

00:01:18.828 --> 00:01:21.664
This is a highly technical field
that requires training in a scientific

00:01:21.664 --> 00:01:24.667
discipline.

00:01:26.211 --> 00:01:28.755
Bioinformaticians can have a range
of educational backgrounds.

00:01:28.755 --> 00:01:32.634
For example, on my team we have folks
who have bachelor's and master's degrees

00:01:32.634 --> 00:01:34.594
in informatics and computer science,

00:01:34.594 --> 00:01:37.597
as well as DJs with postdoc
and work experience in technical fields

00:01:37.597 --> 00:01:40.850
such as biophysics,
medical informatics and biology.

00:01:41.392 --> 00:01:43.061
Even within the field of bioinformatics,

00:01:43.061 --> 00:01:46.106
there are different focuses
between the academic and clinical world.

00:01:46.481 --> 00:01:50.026
Clinical bioinformatics requires a high
level of reproducibility, sensitivity,

00:01:50.026 --> 00:01:53.196
and specificity for consistent processing
from sample to sample.

00:01:53.738 --> 00:01:56.324
This is difficult to conceptualize,
so let's use an analogy to help

00:01:56.324 --> 00:01:58.451
understand the analytical process further.

00:01:58.451 --> 00:02:01.871
Let's use the analogy of reconstructing
the Odyssey by Homer to better understand

00:02:01.871 --> 00:02:05.166
the informatics process needed
for next generation sequencing analysis.

00:02:06.126 --> 00:02:08.545
This book was originally written
in approximately 700

00:02:08.545 --> 00:02:11.548
BC, and has been translated
into English many times.

00:02:12.382 --> 00:02:15.593
Imagine a patient samples
one translation of this epic poem.

00:02:15.885 --> 00:02:18.138
The first step of the process
is equivalent

00:02:18.138 --> 00:02:21.641
to shredding the book into short
multi-character sections, with breakpoints

00:02:21.641 --> 00:02:24.644
randomly occurring within sentences
and even within words.

00:02:25.270 --> 00:02:27.897
Each piece of paper has a variable length
and is then copied

00:02:27.897 --> 00:02:31.192
multiple times, with errors
sprinkled in the text at random.

00:02:31.192 --> 00:02:33.278
Low frequencies.

00:02:33.278 --> 00:02:36.156
The sequencing step is represented
by converting the paper fragments,

00:02:36.156 --> 00:02:38.491
where each letter is translated
into a color code

00:02:38.491 --> 00:02:41.494
for the first 150 characters in each paper
section.

00:02:42.162 --> 00:02:45.373
Typically, a single sequencing
run processes many patient samples

00:02:45.373 --> 00:02:49.127
simultaneously, which would be equivalent
to coding many different translations at

00:02:49.127 --> 00:02:49.961
the same time.

00:02:51.963 --> 00:02:54.591
In our Odyssey analogy,
the first step of the mind for max

00:02:54.591 --> 00:02:55.592
process can be thought

00:02:55.592 --> 00:02:59.470
of as sorting many different translation
of the poem back out from a pile of color

00:02:59.470 --> 00:03:02.682
sequence paper strips
that were indexed by translation version.

00:03:03.766 --> 00:03:05.977
During sequencing,
all translations are color

00:03:05.977 --> 00:03:08.980
coded and decoded into a single document.

00:03:09.189 --> 00:03:11.774
The D multiplexing step sorts out
from a single document.

00:03:11.774 --> 00:03:15.486
The original 150 character segments
that belong to each translation.

00:03:16.821 --> 00:03:18.781
Imagine the reference Odyssey is the very

00:03:18.781 --> 00:03:22.118
first English translation of this poem,
and we are realigning

00:03:22.118 --> 00:03:25.121
a different English translation
back to this first translation.

00:03:26.122 --> 00:03:29.250
There will be different interpretations
for syntax, which could lead to skipping

00:03:29.250 --> 00:03:32.253
words, adding new words,
or even alternate spellings.

00:03:33.463 --> 00:03:37.050
Each 150 character segment of the sample
translation is assigned back to a

00:03:37.050 --> 00:03:41.095
specific page, paragraph, and sentence
in its original reference translation.

00:03:43.014 --> 00:03:46.392
This raw alignment is then polished
using proofreading steps that reduce

00:03:46.392 --> 00:03:49.938
the copies of each fragment, and corrects
errors caused by translation mistakes.

00:03:50.980 --> 00:03:53.024
Proofreading is done by creating an error
model

00:03:53.024 --> 00:03:56.027
based on known words in the Odyssey
at specific locations,

00:03:56.110 --> 00:03:59.113
and these words are used to calibrate
the rest of the decoded characters.

00:04:00.281 --> 00:04:03.284
The metas of my informatics
pipeline is to call variants, or set

00:04:03.284 --> 00:04:05.370
in another way, to identify deviations

00:04:05.370 --> 00:04:08.373
in the sample translation
against the reference translation.

00:04:08.414 --> 00:04:11.626
An example of how this would be done
is to collect overlapping neighboring

00:04:11.626 --> 00:04:14.462
paper strips,
and leveraging the overlapping characters

00:04:14.462 --> 00:04:16.965
to rebuild a larger sentence
or even paragraph.

00:04:16.965 --> 00:04:17.924
They could then, as a whole,

00:04:17.924 --> 00:04:21.427
be realigned back to the reference
translation in order to identify variation

00:04:21.427 --> 00:04:24.514
in characters and words
modified, inserted, or deleted.

00:04:25.765 --> 00:04:27.767
Once the list of deviations are generated,

00:04:27.767 --> 00:04:31.062
the raw locations or in other words, page,
paragraph, and positions

00:04:31.062 --> 00:04:34.065
in the sentence are annotated
with information to give it context.

00:04:35.149 --> 00:04:37.068
These final steps
would allow a medical director,

00:04:37.068 --> 00:04:40.154
acting as the editor,
to review the short list of discrepancies

00:04:40.154 --> 00:04:42.824
between the reviewed translation
against a reference translation

00:04:42.824 --> 00:04:45.243
to understand what changes
occurred before it's published.
