Diversity → design → mutations → application → manufacturing

Design proteins for useful chemistry.

Protein Design Space is an independent scientific webspace for enzyme discovery, protein engineering, biocatalysis, green chemistry, computational design, patent intelligence and scalable biomanufacturing.

PDS
General workflow

From biological diversity to industrial application

The workflow generalises the enzyme-development sequence described in the supplied biocatalysis brochure: identify the enzyme, engineer it, choose an expression system, produce the catalyst, design the reaction, validate the process and translate it to scale. It is expanded here to begin with biological diversity and end with sustainability and IP decisions.

01Diversity & discoveryGenomes, metagenomes, homologues, phylogeny and functional annotation. Search beyond familiar scaffolds before mutating the same protein repeatedly.
02Sequence, structure & mechanismConnect sequence motifs, predicted or experimental structures, active sites, reaction mechanism, channels and conformational constraints.
03Mutation & library designRational substitutions, directed evolution, saturation libraries, second-shell positions, stability networks, interfaces, tunnels and active-learning selection.
04Construct & host selectionChoose expression architecture and host around folding, cofactor needs, soluble expression, secretion, enzyme titre and production economics.
05Expression, production & screeningProduce variants efficiently, purify or use whole-cell/lysate formats, then screen activity, selectivity, stability, solvent tolerance and substrate loading.
06Reaction & catalyst engineeringOptimise pH, temperature, cosolvent, cofactor, donor/acceptor, substrate feed, enzyme loading, immobilisation and catalyst recycling.
07Application & process validationDemonstrate target conversion, product quality, isolation and analytics at proof-of-concept scale, then benchmark gram, kilo and manufacturing-relevant performance.
08TEA, LCA & patent intelligenceTrack cost-of-use, yield, productivity, PMI, E-factor, energy and waste alongside patent claims, sequence scope, legal status and design-around opportunities.
For beginners · AI-enabled protein engineering

Use AI to narrow the search — not to replace experiments

For beginners, AI-enabled protein engineering can be understood as a way to combine evolutionary diversity, protein structure and machine-learning predictions to decide which residues and variants are most worth testing. The goal is to move from thousands or millions of possible mutations to a smaller, information-rich experimental set.

Start with biology
Define the engineering goal

Choose the property you want to improve: activity, selectivity, thermostability, solvent tolerance, expression, substrate range or another measurable function. AI is most useful when the experimental objective is clear.

Diversity
Collect related sequences

Use homologues from databases or metagenomic sources. Sequence diversity reveals what nature has already explored and provides the raw material for conservation and consensus analysis.

Consensus
Align the family

Build a multiple-sequence alignment and look for conserved, variable and consensus residues. Consensus substitutions can be useful starting points, especially for stability, but should be interpreted in structural context.

Structure
Map sequence information onto 3D

Use an experimental or predicted structure to identify catalytic residues, ligand-contacting positions, interfaces, flexible regions and substrate-access tunnels. This separates important functional positions from positions that may tolerate change.

Hotspots
Prioritise positions to mutate

Combine conservation, structure, active-site distance, tunnel analysis and engineering tools such as HotSpot Wizard or FireProt. The output should be a shortlist of positions rather than a random whole-protein library.

AI scoring
Rank candidate substitutions

Protein language models and structure-aware models can score whether substitutions look compatible with natural sequence patterns or a given 3D backbone. Treat these scores as prioritisation signals, not proof of improved function.

AI design
Generate plausible variants

Structure-conditioned design approaches such as ProteinMPNN-style methods can suggest amino acids compatible with a backbone, while ligand-aware approaches can focus design around a substrate or cofactor environment.

Experiment
Build a small informative library

Select a manageable set of variants that represents different hypotheses: consensus changes, active-site changes, tunnel mutations, stability mutations and AI-ranked substitutions. Test the same measurable property across all variants.

Learn & iterate
Feed experimental results back

Use measured activity, stability, expression or selectivity to refine the next round. Even a small dataset can reveal which design assumptions were useful and which regions of sequence space deserve deeper exploration.

A beginner's mental model:
Evolution tells youwhat nature tends to conserve or vary.
Structure tells youwhere residues sit and what they may interact with.
AI helps yourank or generate plausible choices across a huge mutation space.
Experiments decidewhether the variant actually improves the target property.
Beginner worked example

How could you choose one mutation from hundreds of possibilities?

This is a simple hypothetical example to show the logic. Imagine your enzyme has Methionine at position 148 (M148), and you want to know whether this position is worth engineering.

1 · Look at diversity You collect about 500 related sequences and align them.
2 · Check consensus At position 148 you find: Leu 61%, Ile 22%, Met 8%, Val 6%. Your protein's Met is relatively uncommon.
3 · Check the structure M148 is not catalytic, but sits near a hydrophobic pocket and does not make an essential interaction.
4 · Ask AI / design tools A sequence model and a structure-aware design tool both rank Leu and Ile as plausible substitutions.
5 · Test, do not assume You build M148L and M148I, then measure activity, stability and expression against the original enzyme.
What did AI actually do? It did not prove that M148L would be better. It helped reduce a huge mutation space to a small, testable set supported by evolution + structure + model predictions. The experiment still decides whether the variant is useful.
Green chemistry examples

One design logic, many chemistries

Examples below are shown as technology case studies.

Asymmetric amination

Sitagliptin

ω-Transaminase catalysis is a classic example of engineering an enzyme for a demanding chiral-amine transformation.

200 g/Lketone loading in workbook
>99.9% eereported selectivity
90–98%reported yield range
PMI 23vs 30 chemical in workbook

Source DOI · Review DOI

Asymmetric reduction

Montelukast intermediate

A ketoreductase / alcohol-dehydrogenase route demonstrates how protein engineering can replace a sensitive stoichiometric chiral-reduction reagent.

100 g/Lsubstrate loading
>99.9% eecrude product
45°Creported biocatalytic condition
PMI 18–34vs 52 in workbook

Source DOI

Kinetic resolution / hydrolysis

Pregabalin

The lipase-enabled pregabalin manufacturing as an example of improved yield, lower process mass intensity and lower energy demand.

44–45%reported biocatalytic yield
E-factor 17vs 86 chemical
PMI 11.68vs 57.9 chemical
21 MJ/kgvs 118 MJ/kg energy

Source DOI

Platform logic

Beyond APIs

The same workflow applies to industrial enzymes, food biocatalysis, specialty chemicals, environmental enzymes, alternative proteins and precision-fermented products: discover diversity, engineer function, build the expression system, validate application and scale.

Discoverydiversity → candidates
Translationactivity → process value

Protein Design Space is intentionally application-agnostic.

Broader process redesign

Green chemistry is larger than biocatalysis

The specific pharmaceutical process-redesign examples. These are useful benchmarks for waste reduction even when the redesigned route is not necessarily enzyme-based.

APIExample sponsorReported waste decreaseHow it informs protein design
Sertraline HClPfizer / Zoloft92%Shows the scale of process simplification worth targeting.
Sildenafil citratePfizer / Viagra93%Benchmark for solvent, reagent and unit-operation reduction.
CelecoxibPfizer / Celebrex69%Reminds enzyme projects to compare against redesigned chemistry, not legacy chemistry only.
PregabalinPfizer / Lyrica80%Connects catalytic selectivity to large reductions in process waste.
Quinapril HClPfizer / Accupril80%Useful benchmark for route-level sustainability.
SitagliptinMerck / Januvia80%Illustrates how enzyme engineering can become manufacturing innovation.
PaclitaxelBMS / Taxol>90%Highlights the value of route redesign for complex molecules.
NevirapineMedicines for All Institute93%Demonstrates the importance of whole-route optimisation.
GanciclovirRoche Colorado Corp.89%Another benchmark for material-efficiency improvement.
Open scientific toolbox

Resources organised by the protein-design workflow

These resources are grouped by the scientific task they support: learning, sequence discovery, structure, mutation design, channel analysis, molecular visualisation and patent intelligence. Availability and licensing can vary, so check provider terms before commercial use.

Foundations & key reading

Protein engineering concepts from specialist and primary sources

Research overview

Nature — Protein engineering

Research and news collection covering protein-engineering methods and applications.

Open Nature topic ↗
Biocatalysis scale-up

Integrating protein engineering into biocatalytic process scale-up

Perspective linking protein-engineering decisions with industrial biocatalytic process development.

Open DOI ↗
Directed evolution

Protein Design by Directed Evolution

Foundational review covering directed-evolution concepts, library generation and selection for improved protein function.

Open DOI ↗
Semi-rational design

Beyond directed evolution

Review of semi-rational strategies that combine sequence, structure and predictive design.

Open DOI ↗
Enzyme engineering & biocatalysis

Kazlauskas Lab

Academic research resource for enzyme engineering, biocatalysis and protein–function relationships.

Open resource ↗
Sequence discovery & protein properties

Find diversity before designing mutations

Start with homologues, annotations and sequence-level physicochemical profiling before narrowing the design space.

Sequence search

NCBI BLAST

Search homologous protein sequences and expand a known enzyme into a broader diversity set.

Open BLAST ↗
Protein annotation

UniProt

Protein sequences, functional annotations, catalytic information, domains and cross-references.

Open UniProt ↗
Sequence-property profiling

ExPASy ProtScale

Plot amino-acid-scale properties along a protein sequence, including hydrophobicity and other residue-based profiles.

Open ProtScale ↗
Multiple sequence alignment

Clustal Omega

Align homologous protein sequences to identify conserved and variable positions across a protein family.

Open Clustal Omega ↗
Multiple sequence alignment

MAFFT

Fast and flexible multiple-sequence alignment, useful for large and diverse protein families.

Open MAFFT ↗
Multiple sequence alignment

MUSCLE

Multiple-sequence comparison for conservation analysis and mutation-position selection.

Open MUSCLE ↗
Multiple sequence alignment

T-Coffee

Consistency-based sequence alignment with options for integrating different alignment evidence.

Open T-Coffee ↗
Structure prediction & structural references

Move from sequence to three-dimensional hypotheses

Use experimental structures where available, homology models where suitable, and predicted structures to guide engineering hypotheses.

Predicted structures

AlphaFold Protein Structure Database

Search predicted protein structures for structure-guided analysis and comparison.

Open AlphaFold DB ↗
Homology modelling

SWISS-MODEL

Template-based protein structure modelling with model-quality assessment.

Open SWISS-MODEL ↗
Experimental structures

RCSB Protein Data Bank

Experimental macromolecular structures for templates, active-site comparison and structural benchmarking.

Open RCSB PDB ↗
Structure prediction & design

RosettaCommons

Protein modelling and design ecosystem used for structure prediction, docking, redesign and de novo protein design.

Open Rosetta ↗
Structure prediction

I-TASSER

Protein structure and function prediction from sequence using threading and iterative structure assembly.

Open I-TASSER ↗
Remote homology / templates

HHpred

Sensitive profile–profile comparison for remote-homology detection and structure-template discovery.

Open HHpred ↗
Mutation design & stability engineering

Prioritise positions and mutations

Use structure and sequence information to identify candidate engineering positions, then prioritise mutations for stability, activity or altered specificity.

Mutation hotspot discovery

HotSpot Wizard

Web-based protein-engineering workflow for identifying mutagenesis hotspots using structural, functional and evolutionary information.

Open HotSpot Wizard ↗
Stability engineering

FireProt

Protein stability design using structural and evolutionary information to propose stabilising mutations.

Open FireProt ↗
Tunnels, channels & access pathways

Engineer how substrates and products reach the active site

Especially valuable when activity, selectivity or solvent tolerance may depend on access tunnels and internal cavities rather than only catalytic residues.

Tunnel analysis

CAVER

Analyse access tunnels and pathways connecting buried protein regions with bulk solvent.

Open CAVER ↗
Channels & pores

ChannelsDB 2.0 — Methods

Methods and resources for analysing channels, pores and tunnels in biomacromolecular structures.

Open ChannelsDB 2.0 ↗
Molecular visualisation & dynamics

Inspect, compare and communicate structural hypotheses

Use these tools to examine structures, map mutations, view ligands and interfaces, and analyse molecular-dynamics trajectories.

Molecular graphics

PyMOL

Structure visualisation for active sites, residue mapping, ligand inspection and scientific figures.

Open PyMOL ↗
Visualisation & MD analysis

VMD

Molecular visualisation and analysis, widely used with molecular-dynamics trajectories.

Open VMD ↗
Patent landscape & biological-sequence IP

Connect design space with intellectual property space

Use patent searching before and during engineering to understand technology ownership, claimed sequence families and potentially crowded mutation space.

Patent & scholarly landscape

Lens.org

Search patents and scholarly literature, build patent landscapes and investigate technology ownership.

Open Lens ↗
Patent sequence search

Lens PatSeq Finder

Search biological sequences disclosed in patents and connect protein sequence space to patent documents.

Open PatSeq Finder ↗
Founder profile

Dr Vivek Srivastava

Industrial Biotechnology & Biomanufacturing Scientist

PhD molecular and cellular biologist with 19+ years of industrial biotechnology R&D experience spanning industrial enzymes, protein engineering, recombinant proteins and biologics, precision fermentation, downstream processing, analytics, scale-up and commercial translation.

19+years industrial biotechnology R&D
5international PCT patent families in enzyme engineering
1–100 Lpilot-scale process experience
Selected technical impact

Discovery to manufacturing

Industrial enzyme engineering

Protein and enzyme variant discovery, rational design, directed evolution, screening and global application programs.

Precision biomanufacturing

Fermentation, downstream processing, analytics, process troubleshooting and manufacturability.

Alternative proteins

Built an integrated 1–100 L mycoprotein pilot platform and developed a commercial-scale manufacturing concept.

Biologics & recombinant proteins

Multidisciplinary development across therapeutic proteins, peptides, vaccines, antibodies and microbial enzyme programs.

Design principles

What makes an enzyme industrially useful?

Catalytic activity alone is not enough. A useful industrial biocatalyst must survive and perform in the actual process window.

selectivitystabilitysubstrate loadingsolvent tolerancecofactor economyexpression titrereusabilitydownstream compatibilityPMI / E-factor
Patent-aware engineering

Map claims before building libraries

A strong protein-design programme can combine sequence and structure analysis with patent intelligence: identify claimed sequence families and mutation positions, then explore genuinely distinct scaffolds, underused structural regions and alternative process architectures.

Contact

Scientific discussion & collaboration

Protein engineering, enzyme discovery, biocatalysis, green chemistry, patent landscapes, technical reviews and industrial biotechnology.