
During 2025, TAKiR’s website and analysis software became more closely connected to the practical work of running the research. We expanded how people could join programs, provide consent, register DNA kits, upload existing genotype files, receive support, and view genetic connections. At the same time, we developed the analysis system needed to turn laboratory and uploaded data into organized research inputs and results.
That required more than adding screens. A kit had to remain connected to the correct person as it moved through distribution, return, laboratory work, and processing. A consent decision needed a record of the version accepted and a way to withdraw. An analysis result needed to identify its input, method, and processing run. The work of 2025 addressed those connections across the website, database, administrative tools, and analysis service.
Developing program enrollment and consent
In March and April, we built out the program application and consent systems. Program records, eligibility forms, application statuses, administrative notes, and status histories gave the research team a structured way to manage participation. Consent versions and consent records made it possible to associate an individual’s decision with the information presented at the time.
The consent workflow evolved during the year. Work included program-specific consent, clearer instructions, withdrawal, and handling changes to an application’s status. April changes distinguished the Illinois Family Roots program’s consent materials from general website use. June work developed the Family Roots pilot program’s website workflow and waiting-list communications further.
These changes mattered because joining a website, applying to a program, and consenting to research are different actions. The software needed to represent those distinctions and present the appropriate next step to each participant. Administrative views also needed to make the differences visible to the people managing the study.
We revised demographic and profile structures alongside enrollment. The work separated account information, information about the person represented in genetic data, and individual DNA files. Later in the year, a managed-profile structure provided a central person record that could be associated with more than one DNA source.
Following kits through the research process
Kit management became a major area of development. The website gained tools for inventory, registration, assignment, mail requests, sample receipt, laboratory transfer, and genotype receipt. Historical kit information also had to be brought into these workflows so that the system could represent work already underway.
During the summer, we developed barcode recognition and scanning interfaces, bulk registration tools, and ways to associate kit records with users. That work included camera input, uploaded images, text recognition, user searches, and corrections to ambiguous or incomplete historical records. Some approaches were revised or replaced as we worked through the actual administrative tasks.
The November redesign brought kit assignment and lifecycle management into a coordinated interface. Staff could assign a barcode, add tracking, link a kit without an associated user, record a returned sample, mark a kit as sent to the laboratory, or record receipt of genotype data. A mail queue identified requests requiring action, and bulk spreadsheet imports supported selected kit-stage updates.
We also distinguished mailed kits from kits distributed in person. Those paths have different early milestones, even when they eventually converge on laboratory processing. The participant tracker, administrative controls, and email wording were revised to reflect that difference. Statistical summaries were adjusted so stages did not misleadingly count the same kit more than once.
This work turned a collection of kit records into a more usable operational history: what had happened, what was still pending, and what the team needed to do next.
Improving accounts, uploads, and participant support
The account system received repeated attention throughout the year. Work on Cognito integration covered account creation and synchronization, account recovery, authentication errors, and the relationship between the external sign-in service and the website’s own user records. December work addressed optional-phone multifactor authentication, forgotten usernames, password changes, and username case differences during account migration.
DNA upload handling also changed substantially. We developed file validation, metadata storage, profile management, administrative filtering, and processing controls. An April implementation used direct uploads to object storage; in August, that workflow was simplified to direct form submission with consistent size limits, file checks, and filename handling. The goal was a dependable upload process with useful feedback when something failed.
By December, the upload records also distinguished who uploaded a file from the person whose DNA it represented. Source detection replaced reliance on a manually selected testing-company field, and additional changes addressed profile associations, reprocessing, and administrative audit information.
In September, we added a support-ticket system with participant and staff interfaces, conversation responses, attachments, assignment, status tracking, and internal notes. This gave support requests a persistent home within the website instead of leaving their progress implicit in separate exchanges.
The year also included an identity-verification workflow involving document and camera uploads. That system was removed in December. Its development and retirement are both part of the website’s history; it should not be described as a continuing feature of the current site.
Building communications and public research information
Communications work introduced and refined templates, delivery records, subscription preferences, and email groups. Later changes added operational controls, rate limiting, campaign composition, audience selection, and messages tied to kit stages. Moving background work from Celery to Django Q in December also changed how queued and scheduled website tasks were managed.
The public website evolved alongside these systems. Early in the year, we removed the previous Django CMS integration. Subsequent work added a media gallery, revised the homepage and navigation, developed event check-in, and later replaced the Eventbrite-dependent event workflow with native event management. Press coverage and public statistics received their own improvements.
These features supported different forms of engagement: explaining the project, organizing participation, sharing activities and coverage, and answering individual questions. Their underlying models and administrative tools were as important as their public presentation.
Turning laboratory files into analysis inputs
On the analysis side, a major development was the processing path for IDAT files, the raw intensity files produced by the genotyping array. That path required coordinating laboratory batches, sample sheets, array manifests, cluster files, reference sequences, and conversion software.
October and November work developed IDAT-to-genotype conversion and revised the intermediate file handling around BCF, a binary variant format. The pipeline filtered for appropriate variant records, corrected reference-related issues, sorted and indexed files, and produced individual genotype outputs. November work also connected research-kit results to website records and supported raw genotype downloads.
In December, the IDAT processor gained a dedicated container role and service configuration, consolidated reference setup, and direct database integration for exported genotype results. Stage controls allowed work to start or stop at synchronization, conversion, quality control, or export, making it possible to resume selected parts of processing.
The implementation had to address real resource constraints. Changes batched conversion work, managed intermediate files, checked memory, and replaced a memory-intensive validation step with streaming operations. These were concrete engineering responses to processing failures and interrupted runs. They made the workflow more manageable without implying a measured improvement in scientific accuracy.
Organizing the analysis pipeline into stages
The broader batch pipeline developed around quality control and phasing, genetic ancestry analysis, shared-segment detection, and genealogy reconstruction. Each stage needed prerequisites, configuration, logging, output handling, and a way to pass information to the next stage.
Continuous file validation became a separate operational component. During the summer, we developed database polling, container services, incremental processing, and status handling. Later upload work added source detection and common processing support for participant files and curated public datasets, including openSNP and 1000 Genomes.
This also required changes to the website database. Research studies, research participants, research kits, curated samples, and genomic results became explicit records. Analysis-run models and result tables connected quality control and shared-segment output to identifiable processing runs rather than leaving them only as files on a server.
A useful example of the integration work occurred in December. Merging reference samples after phasing could introduce missing, unphased genotypes at positions absent from one input, causing Refined IBD to fail. The batch workflow was changed to merge the samples before joint phasing. Other corrections addressed chromosome naming, sorting and indexing, database constraints, transaction boundaries, and failure propagation when results could not be saved.
These details determine whether a pipeline’s stages can actually exchange data. They are part of the method’s implementation, even when they produce no new page on the website.
Developing shared-segment and family-network research
The year began with substantial work on the Bonsai pedigree reconstruction integration. That work developed preparation of genetic segments and participant metadata, pedigree inference calls, and output handling and visualization.
In September, the lab developed a plan for experiments on using networks of shared DNA to estimate relationships and infer family structures. The planned comparisons would examine how different ways of selecting and grouping samples affect the analysis. The results would help guide updates to the website’s analysis methods and how genetic relationships are presented to participants.
The analysis service also developed integrations for Refined IBD and phasedibd, alongside the earlier IBIS and hap-IBD work, and added RFMix2 execution and input validation. These tools address different questions: shared segments and genetic relationships on one hand, and genetic ancestry along chromosomes on the other. Integrating them required distinct inputs and careful handling of their outputs.
By December, common data structures represented IBD runs, pairs of matching kits, and individual segments. A storage layer translated algorithm-specific output into those records. Corrections made storage failures visible to the surrounding pipeline and made result replacement operate within a database transaction.
Bringing genetic connections into the website
In December, the rebuilt website gained genetic-relative lists, member-profile viewing, research-participant profiles, and interfaces for examining shared segments. User-profile controls distinguished visibility within TAKiR from information shared through external links.
December work also developed statistics for shared-segment analysis and a chromosome display aggregating segments associated with 1000 Genomes reference populations. The public statistics were refined to count unique reference individuals rather than presenting match counts as if they represented different people.
These displays were built on the research models and processing work beneath them. Making a result visible required more than rendering a chart: the site had to identify the analysis run, associate the result with the right person or kit, select the relevant records, and apply the intended access rules.
What 2025 contributed to the present system
During 2025, TAKiR developed the connections between participation and computation. Consent, program enrollment, kit tracking, uploads, support, and research profiles became more structured. Laboratory processing and uploaded-data workflows became more organized, and genetic analysis output gained database representations that the website could use.
The year’s contribution combined software design, research-method integration, operational engineering, and the evaluation of approaches that were sometimes revised or retired. Together, that work moved TAKiR toward the connected participant and analysis system that continued to develop in 2026.
Comments (0)
Be the first to share your thoughts.
Leave a comment