
Following LaKisha David’s 2023 appointment to the faculty at the University of Illinois Urbana-Champaign, 2024 brought a new phase of infrastructure development for The African Kinship Reunion project. TAKiR already had a history of website development and genetic genealogy work. The task in this new phase was to build the infrastructure that would support an expanding research program: a participant website, a durable system for managing research data, and a separate computing environment for genetic analysis.
Much of that work took place beneath the visible website. A page describing the project depends on a web server. A DNA upload depends on storage and a database record that identifies the file. Genetic analysis depends on reference data, compatible file formats, installed software, and a way to follow processing from one step to the next. Updating any of those components requires a deployment process that handles the application and its data together.
During 2024, we worked across those layers. The result was a rebuilt website foundation and a growing analysis environment, supporting our continuing work on preparing genotype data, phasing chromosomes, detecting shared DNA segments, estimating relationships, and inferring family structures from genetic relatedness data.
Establishing the website and its deployment infrastructure
In January 2024, we began rebuilding TAKiR’s website in the current repository. This continued the work of the earlier site, which had already connected DNA uploads to shared-segment results and inferred family trees by December 2021. The rebuild established a new Django application and configuration for Nginx and Gunicorn, the services that receive web requests and run the application. It provided the foundation for expanding the website and its research infrastructure in this next phase.
In May, development concentrated on making deployment repeatable. We worked through container configuration, Python environments, application installation, service startup, static files, and checks that the website could respond after an update. AWS CodeBuild and CodeDeploy became part of the process for preparing and installing the application.
We also built a temporary website testing environment. Deployment scripts launched a test server, assigned its network address, connected it to the appropriate infrastructure, and cleaned it up after use. Browser testing was incorporated into the development workflow, including the practical work of installing a compatible browser and driver on the server.
These tasks were interdependent. A test could not assess the website if the application started with the wrong settings file, the browser could not run, or the server was not reachable. The many corrections to installation paths, environment variables, permissions, and service configuration were the work of making those components function as a system.
By October, infrastructure definitions were organized into CloudFormation templates covering the application, shared infrastructure, and access roles. Keeping these definitions in code made the deployment architecture inspectable and repeatable alongside the website itself.
Updating the database alongside the website
A research website’s database needs to remain aligned with the application that uses it. In late May and June, we developed the machinery for preparing a separate database during deployment, choosing the correct database endpoint, running schema updates, and switching between database environments.
That work continued in November with revisions to the production blue/green deployment approach. In this arrangement, a replacement application environment is prepared alongside the existing one. The implementation coordinated application deployment with database selection, using a stored setting to identify the active database. It also introduced the new.takir.org address for the replacement website.
By the end of November, the database preparation process used a snapshot of the active database to create a writable deployment copy. Related scripts waited for the snapshot and replacement database to become available, updated configuration, and handled cleanup of resources that were no longer needed. Work also covered snapshot export, permissions, deployment logs, and restarting services after a transition.
This was a substantial part of the research infrastructure. The website needed a way to evolve its database and software together as new participant and analysis features were added. Building that process in 2024 established mechanisms that later deployment work would continue to refine.
Building the participant website and its data model
The public website also developed during the year. November and December work added Django CMS integration, blog capability, navigation, and page templates for the project’s mission, research methods, participation, frequently asked questions, privacy, and terms of service.
Those pages created a structure for explaining how genetic analysis and social science research fit within TAKiR. The content system allowed parts of the website to be maintained through an editor, while Django templates defined the surrounding page structure and presentation.
In December, account work added waiting-list and newsletter signup forms and a DNA file upload interface. The upload implementation used private object storage, generated unique file identifiers, and recorded the original filename and testing company in the database. Form hints, navigation, and URL corrections accompanied that work.
The database expanded beyond a user account and an uploaded file. A starter demographics model connected people to DNA profiles. Profiles gained fields for the laboratories and personnel responsible for stages of the work, along with processing status, completion time, and error information. Separate result models provided places to associate individual and batch analysis output with a DNA profile.
These structures made the research workflow more explicit. They distinguished receiving a file from processing it, provided a way to record failure, and linked results back to the profile being analyzed. The account, upload, and demographics components were still being connected and refined at year’s end; they were foundations for the later participant system.
Creating a separate environment for genetic analysis
In November, we established the bagg_analysis repository as part of a separation of concerns: organizing the participant website and genetic analysis as distinct parts of the same research system. The website handled participant interactions, access to data, and presentation of results, while the analysis service took responsibility for data preparation and computational methods. Genetic analysis had already been part of the earlier website; this change gave that work a dedicated service and computing environment.
The separation allowed the website and analysis software to be maintained and updated independently, with their own deployment configurations and software environments. Initial infrastructure definitions and deployment hooks were followed in December by work on installation, background services, logging, and shared storage. The analysis repository included installation scripts for tools used in genotype preparation, phasing, shared-segment detection, ancestry analysis, and pedigree research.
The infrastructure also introduced a background worker for upload events and a batch-processing structure that could look for profiles in the database. These components established the intended connection between uploaded files and analysis. At this stage, parts of the worker’s scientific processing and result-saving logic were placeholders. The substantial analysis work at the end of 2024 was in the research scripts being developed alongside that service infrastructure.
Shared storage was another practical requirement. December changes mounted the analysis file system and corrected access and logging problems so reference files, inputs, and outputs could be used from the computing environment. Installing an algorithm was only one part of the work; the system also had to locate its data, retain its outputs, and provide enough information to diagnose a failed run.
Preparing genotype data and reference materials
Late December work developed a sequence for preparing public genotype data and reference resources. Scripts downloaded and processed openSNP data, parsed genotype files, converted them into Variant Call Format, and merged samples into files suitable for downstream analysis. Supporting routines checked merged files and managed the intermediate files created during conversion.
Reference preparation included downloading 1000 Genomes data, obtaining genetic maps, and selecting reference markers using an Illumina array manifest. These steps addressed a central integration problem: the sample data, reference panel, and analysis tools must agree on the markers and genomic coordinates they use.
The quality-control implementation used PLINK 2 to prepare and filter variants, handle duplicate records, apply configurable missingness and allele-frequency thresholds, and export data by chromosome. It also introduced separate handling for the X and Y chromosomes and mitochondrial DNA, which require different treatment from the numbered autosomes.
The phasing workflow then used Beagle with chromosome-specific reference data and genetic maps. Phasing estimates how a person’s variants are arranged across the chromosome copies inherited from their parents. The scripts sorted and indexed the resulting files and produced chromosome-level statistics, creating organized inputs for the next analysis stage.
Developing shared-segment analysis and experiments for family-tree inference
By the end of December, the analysis code included runners and output-processing routines for IBIS and hap-IBD. Both are established tools for detecting identity-by-descent segments: stretches of DNA inherited from a common ancestor. IBIS operates without requiring phased input, while hap-IBD uses phased genotype data. Our work integrated these tools into the project’s data preparation and result exploration workflow.
The IBIS path converted files into the required format, added genetic-map information, combined chromosome-level results, and examined relationship coefficients and segment properties. The hap-IBD work added routines for running the algorithm, combining and sorting its outputs, and filtering segments by length. This was still an actively developed research workflow, including manual and debugging steps.
Alongside analysis of genetic relatedness data, we developed tools for experimenting with relationship estimation and genetic family-tree inference. The family structures inferred from genetic data were the results to evaluate. Controlled test pedigrees supplied known relationships—a ground truth against which those inferences could be compared.
For this experimental work, we developed a configurable pedigree generator with choices for the number of generations, children, and partners, along with representations of founders and descendants. These deliberately constructed test families, together with pedigree visualization and simulation tools, laid the groundwork for testing whether an analysis could recover a known family structure from genetic data and examining where estimated relationships differed from the ground truth.
What 2024 contributed to the present system
The year’s work connected several kinds of research development. We built web infrastructure that could be deployed and tested, developed database structures for participants and DNA profiles, established a separate analysis service, and implemented research scripts spanning data preparation through shared-segment exploration.
The contribution included selecting and integrating existing scientific software, designing the project’s own data structures and processing routines, and solving the operational problems that arise when those pieces must work together. It gave the next phase of TAKiR a concrete foundation for developing the participant experience and the analysis pipeline further.
The website visible today rests in part on that work: decisions about where data belong, how application changes reach the server, how a file is associated with a person, and how genetic analysis can be organized into inspectable stages.
Comments (0)
Be the first to share your thoughts.
Leave a comment