Sequence & Diversity
Metagenomes, homologues, ancestral reconstruction and scaffold discovery to expand beyond familiar enzyme families.
A scientific webspace for mapping sequence diversity, engineering enzymes, analysing patent landscapes and translating protein designs into robust biocatalysts.
The site is structured around the full path from biological diversity to engineered function, process performance and commercial relevance.
Metagenomes, homologues, ancestral reconstruction and scaffold discovery to expand beyond familiar enzyme families.
Directed evolution, rational design, structure-guided libraries, second-shell residues, tunnels, interfaces and stability engineering.
Protein language models, structure prediction, generative sequence design and active-learning loops for faster search of high-dimensional sequence space.
Enzyme selectivity, substrate loading, solvent tolerance, cofactor economy, immobilisation, reuse and scalable reaction engineering.
Map sequence claims, mutation hotspots, process claims, expiry, legal status and potential design-around or white-space opportunities.
Connect enzyme performance to cost, waste, E-factor, energy, productivity and long-term environmental performance.
A reusable framework for discovering and developing enzymes for new substrates and greener synthesis.
Sitagliptin is a useful industrial case study for how enzyme engineering moves from difficult substrate recognition to high-selectivity manufacturing, and then into immobilisation, catalyst reuse, cofactor economy and new sequence families.
The objective is not merely to catalogue patents, but to translate patent claims into experimentally meaningful design questions.
The most valuable future work may come from combining new biological diversity with modern computational tools and manufacturing-aware selection criteria.
Search uncultured diversity for transaminase scaffolds with novel active-site geometry, stability and solvent tolerance.
Project patent mutation positions onto 3D structures to identify underexplored loops, tunnels, second-shell networks and oligomer interfaces.
Use protein language models and structure-conditioned generation to propose sequences outside familiar natural neighbourhoods.
Close the design-build-test-learn loop by selecting variants that maximise information gain as well as catalytic performance.
Co-optimise enzyme, amino donor, PLP level, immobilisation, solvent and downstream isolation instead of treating them as separate problems.
Prioritise variants that improve process economics and environmental metrics, not merely assay activity.
An independent scientific platform for exploring protein engineering, enzyme discovery, industrial biocatalysis and patent-informed innovation. The site can later be expanded with review articles, interactive patent maps, mutation visualisations, downloadable datasets and project notes.
Independent Researcher · New South Wales, Australia