Updates

TAKiR from December 2021 through 2023: Developing the website and genetic analysis

TAKiR’s current website and analysis system grew out of earlier work on connecting genetic data, inferred family relationships, and a website that people could use to examine them. From December 2021 through 2023, much of that work was recorded in the starlight repository, followed by a separate website experiment in starbrightdir.

By this period, the project was already underway. December 2021 is the starting point for this account, not the beginning of TAKiR.

LaKisha David graduated from the University of Illinois Urbana-Champaign (UIUC) in May 2021 with a PhD in Human Development and Family Studies. From 2021 to 2023, she held a postdoctoral appointment in the Ethical, Legal and Social Implications (ELSI) of Genetics and Genomics postdoctoral training program at the University of Pennsylvania’s Perelman School of Medicine. In August 2023, she began her faculty position in UIUC’s Department of Anthropology. The website and analysis work described here spans her postdoctoral period and the beginning of her faculty appointment.

Connecting people, uploaded files, and genetic data

In December 2021, the earlier Django website developed a connected path from DNA file upload toward genetic analysis. Its data model distinguished a website user, a participant represented in the research, and a DNA profile containing an uploaded genotype file.

Participants had a genotype identifier and fields used by the analysis, including birth year and chromosome sex. Uploaded profiles recorded their testing company and file status. Those distinctions allowed the system to associate a file with the person whose DNA it represented and to retain information needed for later relationship inference.

The implementation also needed to handle the practical differences among genotype files. Supporting code used the Lineage library to load data, check the genome build, and remap data when needed. Preparation scripts converted and merged genotype files, then used PLINK to create inputs for shared-segment analysis.

This was already work on a research workflow, not simply a collection of informational pages. The website and its processing scripts had to agree on participant identifiers, file locations, input formats, and the status of the data being analyzed.

Preparing data for shared-segment detection

The December 2021 processing code included quality-control steps for genetic variants and missing measurements. The active PLINK command selected nucleotide variants, excluded duplicate identifiers, limited records to two alleles, and applied filtering settings before writing the files needed for analysis.

These settings were part of an evolving implementation. The source retained earlier alternatives and notes about problems encountered with filtering, showing the practical work of adapting a pipeline to its inputs. Preparing a file in the right format did not guarantee that its markers, sample coverage, and reference coordinates were appropriate for the next tool.

The shared-segment stage used IBIS, an established method for identifying identity-by-descent segments. Such segments represent stretches of DNA inherited from a common ancestor. The implementation added genetic-map information to the prepared files, ran the detector, and brought its output back into the application.

The database represented more than a single match score. Segment records included chromosome, physical start and end positions, genetic positions and length, marker counts, and available error measures. This preserved information needed to inspect where two people shared DNA and how a segment had been reported.

Integrating pedigree inference with the website

Bonsai integration was a major part of the work in December 2021. Bonsai is an established pedigree-inference tool; the project’s work was to prepare its inputs, run it in the context of the website’s participant records, and organize its output for examination.

The integration assembled genetic segments and participant information for a selected focal person. Birth year supplied an approximate age, and the pipeline translated participant data into the format expected by the inference tool. Results were stored as pedigree relationships and nodes associated with the participant.

That connection required considerable result-handling work. Shared-segment summaries, inferred ancestor information, and tree structures had to be represented in the database and translated into displays. By December 19, 2021, the repository included a basic SVG pedigree display. Further changes through the end of December 2021 developed the diagram generation, result loading, and overall Bonsai workflow.

The code also explored how to handle the time spent loading and rendering results. It used concurrent workers for selected database and diagram tasks while retaining a separate loop for pedigree inference. These were practical implementation decisions around a research algorithm, intended to make its output usable through the website.

Showing the evidence behind a match

The earlier result pages exposed several levels of detail. A participant could view a list of matches and move from a summary of shared DNA into individual segments. The match table included total shared DNA, segment count, the largest segment, and inferred information about shared ancestors.

January 2022 work added a connecting path between two people’s genotype identifiers. That path could be followed in the inferred tree, linking the match table to the pedigree diagram. It was an early attempt to help readers move between a pairwise result and the larger family structure proposed by the analysis.

The page also highlighted participant records indicating that all four grandparents were born in Africa. This was a distinction based on information associated with the participant record, not a claim that the highlighting itself inferred a person’s ancestry from DNA.

These details matter to the history of the current website. The questions behind today’s relative lists, chromosome displays, and family-tree views were already present: what evidence supports a connection, how can that evidence be examined, and how should an inferred relationship be presented to a reader?

Managing processing and explaining incomplete results

The early pipeline also needed a way to distinguish new work from previously processed data. Its orchestration checked for DNA profiles updated since the last run and recorded when a new run began. It could select updated records or accept a researcher-specified list of genotype identifiers for analysis.

Processing updated participant status based on whether the uploaded data passed the initial checks. The result page then distinguished several situations: no uploaded file, a file waiting for the next run, a file that failed a check, and a completed check without matches meeting the current criteria.

That distinction is important. “No result” can mean that analysis has not happened, that a file could not be used, or that the analysis did not identify a qualifying match. The December 2021 and early January 2022 work began representing those possibilities explicitly instead of displaying the same empty page for all of them.

Instructions for site users were added alongside the analysis and display changes. The project was developing both the processing machinery and the explanations needed for people to understand its state.

Maintaining and revising the earlier website

During 2022, work continued on the website’s presentation and operation. October 2022 revisions reorganized the homepage, about and personnel pages, banners, navigation, and related layouts. Registration fixes followed, alongside updates to dependencies used by the Django application and its visual components.

Maintenance continued into early 2023. Changes addressed form protection, login and contact-page errors, and dependency updates. These tasks supported the existing site even when they did not add a new research method or a major new interface.

In April 2023, the site’s signup, contact, and email-related functions were removed or disabled. That is part of the history as well: the earlier website did not simply expand continuously until the current version replaced it. Its available functions changed as the project’s website approach evolved.

Exploring a different website structure in 2023

In September 2023, starbrightdir recorded a separate Django CMS experiment. It introduced a CMS-based site structure, page templates with editable placeholders, shared navigation and footer elements, and blog integration.

The experiment explored a different division of responsibility between code and content. Templates defined the surrounding layout, while CMS placeholders provided places for content maintained through an editor. Deployment configuration also went through several Elastic Beanstalk revisions during that short period.

This was an exploratory branch of the website’s history. It should not be mistaken for the creation of the present bagg_website repository or for a completed migration of every earlier research feature. It records work on how a future website might be structured and maintained.

What this period contributed to the later project

By the time the current website repository began in 2024, TAKiR already had experience connecting participant records, uploaded DNA files, shared-segment analysis, pedigree inference, and web presentation.

That experience included building and maintaining a research website on Amazon Web Services (AWS), with Amazon RDS databases, network configuration, application deployment, and ongoing operation. The project had already worked through the connections between cloud infrastructure and the application it supported: database access, file conversion, result loading, processing status, and guidance for site users. The 2024 rebuild drew on this existing experience with both genetic analysis and AWS infrastructure.

The later system would revise many of those implementations and separate website and analysis infrastructure more clearly. The project also expanded beyond uploaded genotype files to collecting saliva samples and carrying out our own processing of the laboratory’s raw array data. That work extends from IDAT intensity files through conversion into genotype data in Variant Call Format (VCF), quality control, and downstream genetic analysis. The earlier website had already analyzed uploaded genotypes; the later work brought more of the preceding data-preparation process into the project’s own system.

Documenting December 2021 through 2023 makes that continuity visible. The 2024 infrastructure rebuild was a new phase of an existing project, built on work that had already connected genetic analysis to a participant-facing website.

Comments (0)

Be the first to share your thoughts.

Leave a comment

Markdown supported: **bold**, *italic*, `code`, [link](url)
Sign in for a verified badge

Back to the blog