Logo Sequensolutions
Research Article

Immunoinformatics-Driven Design and In Silico Validation of a Multi-Epitope Subunit Vaccine Targeting Norovirus

Nitish Kumar R1* Kesiya Joy1 Parvana Nair1

1 Department of Life Sciences, Garden City University, Bengaluru, India

* email: nitishkumar271120@gmail.com

Published: Jan 2026 Category: Immunoinformatics Preprint DOI: 10.21203/rs.3.rs-8606904/v1

Abstract

Norovirus, a non-enveloped, positive-sense single-stranded RNA virus belonging to the Caliciviridae family, is a major causative agent of acute gastroenteritis (AGE) outbreaks worldwide. It is primarily transmitted via the fecal–oral route, with clinical manifestations including abdominal pain, watery diarrhoea, nausea, and vomiting. In the United States alone, norovirus is estimated to cause approximately 19–21 million cases annually. Emerging variants, such as the GII.17 genotype, have been implicated in an increasing number of outbreaks across multiple countries, highlighting the urgent need for effective preventive strategies. In this study, an immunoinformatics-based approach was employed to design and evaluate a multi-epitope subunit vaccine candidate against norovirus. Three viral proteins—capsid protein (UniProt ID: A7YK10), small protein (A7YK11), and polyprotein (A7YK09)—were selected for epitope-based vaccine design. B-cell and T-cell epitopes were predicted using the Immune Epitope Database (IEDB) and subsequently screened for antigenicity (VaxiJen), allergenicity (AllerTOP v2.1), toxicity (ToxinPred), and population coverage. Physicochemical properties were evaluated using ProtParam, and secondary structure analysis was performed to assess the structural feasibility of the vaccine construct. The finalized multi-epitope vaccine construct was further subjected to molecular docking analyses to evaluate its binding affinity with key immune receptors, including MHC class I, MHC class II, and Toll-like receptor 4 (TLR4), providing insight into its potential immunogenic interactions. While the computational analyses indicate that the designed construct is a promising vaccine candidate, further experimental validation, including in vitro expression and in vivo immunogenicity and efficacy studies, will be required to confirm its protective potential.

Keywords: Norovirus; acute gastroenteritis; multi-epitope vaccine; immunoinformatics; molecular docking; immunogenicity

1. Introduction

Norovirus (NoV), a non-enveloped, positive sense single-stranded RNA virus classified under the Caliciviridae family, is globally recognized as the leading cause of non-bacterial acute gastroenteritis (AGE). This highly contagious pathogen affects individuals across all age groups and significantly contributes to the worldwide burden of gastrointestinal illnesses and associated mortality (Lanata et al., 2013). Clinical manifestations of Norovirus infection commonly include abdominal pain, watery diarrhoea, nausea, and vomiting, with transmission primarily occurring via the fecal-oral route (Lew et al., 2012).

The scale of its impact is substantial; in the United States alone, Norovirus is estimated to be responsible for approximately 19 to 21 million cases annually. The virus poses a particularly serious threat to vulnerable populations, such as infants and the elderly, largely due to its remarkable ease of transmission and its resilience in diverse environmental conditions (Ahmed et al., 2014; Pires et al., 2015).

The dynamic nature of Norovirus, exemplified by the recent implication of emerging variants like the GII.17 genotype in an increased number of outbreaks across multiple countries, presents a continuous challenge to public health (Afgan et al., 2018). This constant evolution of viral strains directly implies a significant hurdle for traditional vaccine approaches, as antigenic drift can rapidly reduce the efficacy of vaccines designed against specific, static targets. While in silico predictions offer a powerful and rapid initial step in vaccine design, their ultimate utility in addressing the dynamic nature of viral evolution necessitates rigorous experimental validation (Somana et al., 2020).

Despite numerous international efforts, a licensed Norovirus vaccine or specific antiviral therapy remains unavailable. This critical gap in treatment and prevention is primarily a consequence of several formidable biological and technical hurdles (Tanaka et al., 2013). The virus exhibits extensive genetic variability, particularly prominent in genogroups GI and GII, with the GII.4 genotype being notably the most common and virulent. This high mutation rate, leading to frequent antigenic variations, is a major factor contributing to recurring global outbreaks and renders traditional immunization strategies, which often target single or limited strains, largely inadequate (Patel et al., 2021).

The pivot to computational methods, therefore, becomes not merely an alternative but a strategic imperative to circumvent these long-standing obstacles, enabling the identification of conserved sequences that can provide broader protection against evolving strains and accelerate the vaccine discovery pipeline. In response to the challenges posed by Norovirus, computational vaccinology and immunoinformatics have emerged as innovative and cost-effective methods for creating multi-epitope vaccines (Atmar et al., 2020). These advanced approaches leverage bioinformatics tools and algorithms to identify conserved sequences across various viral strains, enabling the rational design of vaccine candidates that can elicit robust and lasting immune responses efficiently.

2. Materials and Methods

2.1 Target Protein Sequence Retrieval and Selection

Norovirus protein sequences were retrieved through a systematic database search. The National Center for Biotechnology Information (NCBI) database was used to identify relevant Norovirus genomic information. Based on extensive literature review and antigenic relevance, three viral proteins were selected for vaccine design: the capsid protein, small protein, and polyprotein. Protein sequences corresponding to the Murine norovirus GV/WU24/2005/USA strain (OX = 463715) were obtained in FASTA format from the UniProt database. The selected proteins and their respective UniProt identifiers were A7YK10 (capsid protein), A7YK11 (small protein), and A7YK09 (polyprotein).

2.2 Overall Workflow for Epitope-Based Vaccine Design

The selected Norovirus protein sequences were subjected to a comprehensive immunoinformatics pipeline involving epitope prediction, screening, vaccine construct design, structural modeling, molecular docking, immune simulation, and codon optimization. A schematic representation of the overall workflow is shown in Figure 1.

Figure 1
Figure 1. Schematic representation of the steps involved in epitope-based vaccine designing.

2.3 B-Cell Epitope Prediction

Linear B-cell epitopes were predicted from the selected protein sequences using the Immune Epitope Database (IEDB) resource. Predicted epitopes were evaluated for their antigenic potential and suitability for inclusion in the vaccine construct.

2.4 T-Cell Epitope Prediction

T-cell epitope prediction was performed using IEDB tools to identify peptides capable of binding to major histocompatibility complex (MHC) molecules. Both MHC class I–restricted epitopes (for CD8⁺ T cells) and MHC class II–restricted epitopes (for CD4⁺ T cells) were predicted to ensure the induction of cell-mediated immune responses.

2.5 Cytotoxic T Lymphocyte (CTL) Epitope Prediction

CTL epitopes binding to MHC class I molecules were predicted using the IEDB MHC-I binding tool. Predictions were based on peptide binding affinity, proteasomal C-terminal cleavage efficiency, and transporter associated with antigen processing (TAP) transport potential.

2.6 Helper T Lymphocyte (HTL) Epitope Prediction

Helper T lymphocyte (HTL) epitopes were predicted using the IEDB MHC-II binding tool. Epitopes with strong binding affinity to commonly occurring human leukocyte antigen (HLA) class II alleles were prioritized to enhance CD4⁺ T-cell–mediated immune responses.

2.7 Population Coverage Analysis

Population coverage analysis was conducted using the IEDB-AR v2.22 population coverage tool to evaluate the global applicability of the selected CTL and HTL epitopes.

2.8 Vaccine Construct Design

A multi-epitope subunit vaccine (MEV) construct was designed by assembling the selected CTL and HTL epitopes using appropriate linker sequences. Epitope selection was based on high antigenicity, non-allergenicity, non-toxicity, conservation across strains, strong HLA binding affinity, and broad population coverage. The construct was designed to optimize immunogenicity, stability, and safety.

2.9 Physicochemical Property Analysis

The physicochemical properties of the designed vaccine construct, including molecular weight, theoretical isoelectric point (pI), instability index, aliphatic index, and grand average of hydropathicity (GRAVY), were evaluated using the ProtParam tool.

2.10 Antigenicity and Solubility Prediction

Antigenicity of the vaccine construct was assessed using the VaxiJen v2.0 server with a threshold value of 0.4. Protein solubility upon recombinant expression in Escherichia coli was predicted using the SOLUPROT web server.

2.11 Secondary Structure Prediction

Secondary structure prediction of the vaccine construct was performed using the PSIPRED server. The predicted proportions of alpha helices, beta strands, and random coils were analyzed.

2.12 Tertiary Structure Prediction and Validation

Tertiary structure prediction was performed using the Robetta web server. Structural validation of the predicted model was conducted using PROCHECK via the SAVES v6.0 server. Ramachandran plot analysis was used to assess stereochemical quality, and the ProSA-web server was employed to evaluate overall model quality and structural stability.

2.13 Molecular Docking Analysis

Molecular docking was performed to evaluate interactions between the vaccine construct and key immune receptors, including Toll-like receptor 4 (TLR4), MHC class I, and MHC class II molecules. Docking simulations were carried out using the HADDOCK v2.2 web server.

2.14 In Silico Immune Simulation

In silico immune simulation was performed using the C-ImmSim server to predict immune responses elicited by the vaccine construct. Simulations were conducted at time steps 1, 84, and 168, corresponding to three vaccine doses administered at four-week intervals.

2.15 Codon Optimization and In Silico Cloning

Codon optimization of the vaccine construct was performed using the Java Codon Adaptation Tool (JCAT) to enhance expression efficiency in E. coli. The optimized nucleotide sequence was subsequently cloned in silico into the pET-28a(+) expression vector using SnapGene software.

3. Results

3.1 Protein Sequence Retrieval

The target Norovirus protein sequences were successfully retrieved from the UniProt database. Three proteins were selected for epitope prediction and vaccine design: the capsid protein (UniProt ID: A7YK10), small protein (UniProt ID: A7YK11), and polyprotein (UniProt ID: A7YK09).

3.2 Epitope Screening and Selection

The capsid protein, small protein, and polyprotein sequences were submitted to the Immune Epitope Database (IEDB) server for the prediction of B-cell, cytotoxic T lymphocyte (CTL), and helper T lymphocyte (HTL) epitopes. Following comprehensive screening, a total of six B-cell epitopes, six MHC class I (CTL) epitopes, and nine MHC class II (HTL) epitopes were selected.

Tables 1–3 summarize the selected B-cell, MHC class I, and MHC class II epitopes along with their respective immunological properties.

Table 1. Selected B-cell epitopes
SEQUENCE TOXICITY ALLERGENICITY ANTIGENICITY VAXIJEN SCORE
HVNGTLLGTTPVSGSWVS NON TOXIN NON-ALLERGEN ANTIGEN 0.5223
PLDLVDGRVRAVPRSVYFFQDVLPEYNDGLL NON TOXIN NON-ALLERGEN ANTIGEN 0.6936
SEDEVNPALL NON TOXIN NON-ALLERGEN ANTIGEN 0.8404
DELVPKQDEKYQK NON TOXIN NON-ALLERGEN ANTIGEN 0.5609
TKGPHPGKPELTPLGA NON TOXIN NON-ALLERGEN ANTIGEN 1.1375
FGTMDAEPTQERSA NON TOXIN NON-ALLERGEN ANTIGEN 0.9676
Table 2. Selected MHC class I (CTL) epitopes
SEQUENCE TOXICITY IFN GAMMA ALLERGEN ANTIGEN VAXIJEN
AVDWSGTRYY NON TOXIN POSITIVE NON-ALLERGEN ANTIGEN 0.4291
LPSLRGGSW NON TOXIN POSITIVE NON-ALLERGEN ANTIGEN 1.0971
SWVPRLFQL NON TOXIN POSITIVE NON-ALLERGEN ANTIGEN 0.5009
SEDPVPALL NON TOXIN POSITIVE NON-ALLERGEN ANTIGEN 0.4575
ETLPGHAQR NON TOXIN POSITIVE NON-ALLERGEN ANTIGEN 0.4575
SESEDEVNY NON TOXIN POSITIVE NON-ALLERGEN ANTIGEN 0.5439
Table 3. Selected MHC class II (HTL) epitopes
SEQUENCE TOXICITY IFN GAMMA ALLERGEN ANTIGEN VAXIJEN
DIEMLGAQVQAQAQA NON-TOXIN POSITIVE NON ALLERGEN ANTIGEN 0.9669
LGAQVQAQAQAQENA NON-TOXIN POSITIVE NON ALLERGEN ANTIGEN 0.9256
KHDIEMLGAQVQAQA NON-TOXIN POSITIVE NON ALLERGEN ANTIGEN 0.7088
DQAPYQGKVYASLAA NON-TOXIN POSITIVE NON ALLERGEN ANTIGEN 0.4022
DFNFVYLTPPIERTV NON-TOXIN POSITIVE NON ALLERGEN ANTIGEN 1.3365
AQDWNVDPQPFIPS NON-TOXIN POSITIVE NON ALLERGEN ANTIGEN 1.0250
LKRYGLLPTRADKEE NON-TOXIN POSITIVE NON ALLERGEN ANTIGEN 0.9887
VTAFKAMAADAGIPW NON-TOXIN POSITIVE NON ALLERGEN ANTIGEN 0.4609
KRYGLLPTRADKEEG NON-TOXIN POSITIVE NON ALLERGEN ANTIGEN 1.2715

3.3 Population Coverage Analysis

Population coverage analysis was performed using the IEDB Population Coverage Tool. The combined set of six CTL and nine HTL epitopes demonstrated an overall worldwide population coverage of 98.96%, indicating that the selected epitopes have the potential to elicit immune responses in a large proportion of the global population.

3.4 Vaccine Construct Design

The final multi-epitope vaccine construct was designed by assembling the six selected CTL epitopes and nine HTL epitopes. β-defensin 3, a Toll-like receptor 4 (TLR4) agonist, was incorporated at the N-terminal region as an adjuvant to enhance immune stimulation.

Figure 3
Figure 3. Illustration of overall architecture of the multi-epitope vaccine construct.

3.5 Physicochemical Properties and Solubility

The physicochemical properties of the final vaccine construct were analyzed using the ProtParam tool. The construct consisted of 433 amino acids with a molecular weight of 46,095.07 Da and a theoretical isoelectric point (pI) of 8.88. The instability index was calculated to be 29.20, indicating that the protein is predicted to be stable. Protein solubility analysis using the SOLproT server yielded a solubility score of 0.894, indicating that the construct is highly soluble.

3.6 Secondary Structure Prediction

Secondary structure analysis of the vaccine construct was performed using the PSIPRED server. The results showed that 71 amino acids (16.40%) formed alpha helices, while 60 amino acids (13.86%) contributed to beta strands. The majority of the sequence comprised random coils (302 amino acids; 69.7%).

Figure 4
Figure 4. Depiction of predicted secondary structure of the refined vaccine construct.

3.7 Tertiary Structure Prediction and Validation

The tertiary structure of the vaccine construct was predicted using the Robetta server. Model 1 was selected and refined using the GalaxyRefine server. Structural validation revealed that 88.9% of residues were located in the most favored regions of the Ramachandran plot. The refined model achieved a MolProbity score of 1.55 and a QMEAN Z-score of −0.97.

Figure 5a Figure 5b
Figure 5. Validation of the 3D final vaccine model. (A) Ramachandran analysis. (B) QMEANDIsco 3D structure validation.

3.8 Molecular Docking and Protein–Protein Interactions

Protein–protein docking analysis was conducted to evaluate interactions between the vaccine construct and key immune receptors, including TLR4, MHC class I, and MHC class II molecules. The vaccine–TLR4 complex exhibited the most favorable interaction, with a HADDOCK score of −103.9 ± 11.0.

Figure 6a Figure 6b Figure 6c
Figure 6. Illustration the docked complexes with (A) MHC-I, (B) MHC-II, and (C) TLR4.

3.9 In Silico Immune Simulation Analysis

Immune simulation using the C-ImmSim server predicted robust immune responses following administration of the vaccine construct. The simulation demonstrated a substantial increase in antibody levels and populations of active B cells, helper T cells, and cytotoxic T cells.

Figure 7
Figure 7. Simulation of immune responses in three doses administration regimen.

3.10 Molecular Dynamics Simulation

Molecular dynamics (MD) simulation was performed to assess the stability of the vaccine–TLR4 complex over a 100 ns simulation period. The root mean square deviation (RMSD) values remained stable within the range of 1–1.5 Å.

Figure 8a Figure 8b
Figure 8. (A) Molecular dynamics simulation RMSD. (B) RMSF of the complex.

3.11 Codon Optimization and In Silico Cloning

Codon optimization of the vaccine construct was performed using the JCAT server. The optimized nucleotide sequence comprised 1395 nucleotides, with a GC content of 56.99% and a codon adaptation index (CAI) of 1.0.

The optimized vaccine gene was cloned in silico into the pET-28a(+) expression vector between the XhoI and NdeI restriction sites using SnapGene software.

Figure 9
Figure 9. In silico restriction cloning of the final vaccine construct into pET28a(+) expression vector.

4. Discussion

The increasing global burden of Norovirus infections, coupled with the continued absence of a licensed vaccine, underscores the urgent need for alternative and efficient vaccine development strategies (Atmar & Estes, 2020; Patel et al., 2009). In this study, an immunoinformatics-driven framework was employed to design and evaluate a multi-epitope subunit vaccine candidate targeting Norovirus. By integrating conserved antigenic regions from the capsid protein (A7YK10), small protein (A7YK11), and polyprotein (A7YK09), the design strategy aimed to address the extensive genetic diversity and antigenic variability characteristic of Norovirus.

The epitope selection process was guided by stringent immunological and safety criteria, including antigenicity, non-allergenicity, non-toxicity, IFN-γ induction potential, and broad HLA population coverage. Physicochemical characterization of the vaccine construct revealed favorable properties for stability and expression.

Structural analysis further reinforced the robustness of the designed construct. Secondary structure prediction indicated a predominance of flexible coil regions, which may facilitate effective epitope exposure and immune recognition. Molecular docking and molecular dynamics simulations provided mechanistic insights into the interaction between the vaccine construct and key immune receptors. The in silico immune simulation results complemented the structural and docking analyses by predicting robust humoral and cellular immune responses.

Despite these promising results, it is important to acknowledge the limitations inherent to computational studies. Therefore, experimental validation remains essential. Future studies should focus on in vitro expression and purification of the vaccine construct, followed by immunogenicity assessment.

5. Conclusion

In this study, an immunoinformatics and reverse vaccinology approach was successfully applied to design a rational multi-epitope subunit vaccine candidate against Norovirus. Comprehensive in silico analyses demonstrated that the final vaccine construct possesses favorable physicochemical properties, structural stability, and strong interactions with key immune receptors. Overall, these findings suggest that the designed multi-epitope vaccine construct represents a promising computational candidate for Norovirus vaccine development.

Acknowledgement

The authors extend their gratitude to Dr. Ramchandra Prasad and Dr. Shanmuga Priya for their valuable guidance and support.

Disclosure statement

No potential conflict of interest was reported by the authors.

References

  1. Afgan, E., et al. (2018). The Galaxy platform for accessible, reproducible and collaborative biomedical analyses: 2018 update. Nucleic Acids Research.
  2. Ahmad, I., et al. (2020). Development of multi-epitope subunit vaccine for protection against norovirus infections based on computational vaccinology. Journal of Biomolecular Structure and Dynamics.
  3. Ahmed, S. R., & Abid, M. (2022). Structural vaccinology and immunoinformatics approach for designing a multi-epitope vaccine against norovirus. Molecular Immunology.
  4. Atmar, R. L., & Estes, M. K. (2020). Norovirus vaccine development: Next steps. Expert Review of Vaccines.
  5. Azim, K. F., et al. (2020). Immunoinformatics approaches for designing a novel multi-epitope peptide vaccine against human norovirus. Informatics in Medicine Unlocked.
  6. Bartsch, S. M., et al. (2018). The clinical and economic burden of norovirus gastroenteritis in the United States. Journal of Infectious Diseases.
  7. Cannon, J. L., & Vinjé, J. (2020). Evolution and genetic classification of noroviruses. ASM Press.
  8. Chen, J., et al. (2024). Advances in human norovirus research: Vaccines, genotype distribution and antiviral strategies. Computers in Biology and Medicine.
  9. Chhabra, P., et al. (2019). Updated classification of norovirus genogroups and genotypes. Journal of General Virology.
  10. Das, S., et al. (2021). Limitations and challenges in computational vaccine design. Briefings in Bioinformatics.
  11. Doytchinova, I. A., & Flower, D. R. (2007). VaxiJen: A server for prediction of protective antigens. BMC Bioinformatics.
  12. Gasteiger, E., et al. (2005). Protein identification and analysis tools on the ExPASy server. The Proteomics Protocols Handbook.
  13. Grote, A., et al. (2005). JCat: A novel tool to adapt codon usage of a target gene. Nucleic Acids Research.
  14. Jones, D. T. (1999). Protein secondary structure prediction based on position-specific scoring matrices. Journal of Molecular Biology.
  15. Lo, Y. C., et al. (2023). The changing landscape of norovirus genotypes. Journal of Medical Virology.
  16. Magnan, C. N., et al. (2009). SOLpro: Accurate sequence-based prediction of protein solubility. Bioinformatics.
  17. Matos, A. O., et al. (2023). Immunoinformatics-guided design of a multi-valent vaccine against rotavirus and norovirus. Computers in Biology and Medicine.
  18. Patel, M. M., et al. (2009). Noroviruses: A comprehensive review. Journal of Clinical Virology.
  19. Tan, M., & Jiang, X. (2014). Histo-blood group antigens: A common niche for norovirus and rotavirus. Expert Reviews in Molecular Medicine.
  20. van Zundert, G. C. P., et al. (2016). The HADDOCK2.2 web server: User-friendly integrative modeling of biomolecular complexes. Journal of Molecular Biology.
  21. Vita, R., et al. (2019). The Immune Epitope Database (IEDB): 2018 update. Nucleic Acids Research.