Bioinformatics · Artificial Intelligence · Life Sciences
Aditya BioNova AnalyticsGen AI · Artificial Intelligence · Data Science · Bioinformatics
Founder · Scientific Leadership

Dr. Sambhaji Balaso Thakar

Senior Bioinformatician · Data Scientist
MSc BioinformaticsPh.D. BiotechnologyPost-Doctoral Fellow, CAU, Seoul, South KoreaPost-Doctoral Fellow, Anhui Agricultural University (AAU), Hefei, China
Dr. Sambhaji Balaso Thakar
FounderAditya BioNova Analytics
Post-Doctoral Research

Chung-Ang University (CAU), Seoul, South Korea · Anhui Agricultural University (AAU), Hefei, China

Professional Roles

Senior Bioinformatician · Data Scientist

Education

MSc Bioinformatics, S.R.T.M. University, Nanded · PhD Biotechnology (Structural Bioinformatics), Shivaji University, Kolhapur

Profile

Executive Summary

Dr. Sambhaji Balaso Thakar is a PhD-qualified Bioinformatician and Data Scientist with international postdoctoral research experience at Chung-Ang University (CAU), Seoul, South Korea and Anhui Agricultural University (AAU), Hefei, China.

His expertise spans artificial intelligence, machine learning and large-scale biological data analytics, with a focus on biological database development, next-generation sequencing (NGS), structural bioinformatics, systems biology and retrosynthesis models.

His peer-reviewed research contributions span cancer-therapy reviews, transcriptomics, microbiome analysis and biological databases.

He completed postdoctoral research at Chung-Ang University in Seoul, South Korea (2020–2021), and Anhui Agricultural University in Hefei, China (2018–2020).

His experience also includes bioinformatics teaching, scientific data curation and research collaborations.

At Aditya BioNova Analytics, he leads scientific project scoping and review, with an emphasis on transparent methods and reproducible analysis.

Skills

Core Competencies

01Artificial Intelligence & Machine Learning

  • Support Vector Machines (SVM)
  • Decision Trees
  • K-Nearest Neighbors (KNN)
  • Bayesian models
  • K-Means clustering
  • Synthetic accessibility and retrosynthesis models
  • Data analytics

02Bioinformatics & Computational Biology

  • Next-generation sequencing (NGS) analysis
  • Sequence and structural analysis
  • Phylogenetic analysis
  • Biological database development
  • Scientific data curation

03Bioinformatics Tools & Omics

  • Multi-omics (Genomics, Transcriptomics, Proteomics, Metabolomics)
  • Sequence alignment (BLAST, Bowtie2)
  • Variant calling (GATK)
  • Structural biology (AlphaFold, PyMOL)
  • Nextflow
  • Snakemake

04Programming & Core Languages

  • Python
  • R
  • SQL
  • C++
  • Java
  • Bash / Linux
  • Pandas
  • NumPy

05Data Science & Machine Learning

  • Predictive modeling
  • Feature engineering
  • Statistical modeling
  • Exploratory data analysis (EDA)
  • Scikit-learn
  • XGBoost
  • Dimensionality reduction (PCA, t-SNE, UMAP)

06Deep Learning & NLP

  • TensorFlow
  • PyTorch
  • Neural networks
  • Natural Language Processing (NLP)

07Generative AI & LLMs

  • Large Language Models (LLMs)
  • Retrieval-Augmented Generation (RAG)
  • LangChain
  • LlamaIndex
  • Prompt engineering
  • Fine-tuning (LoRA, QLoRA)
  • BioBERT
  • DNABERT
  • OpenAI GPT-4
  • Google Gemini
  • Claude
  • Llama 3
  • Hugging Face Transformers

08Vector Databases, MLOps & Cloud

  • Pinecone
  • ChromaDB
  • Milvus
  • Docker
  • Kubernetes
  • AWS (Bedrock / Lambda)
  • GCP
  • Git
  • High-Performance Computing (HPC)

09Data Visualization

  • Power BI
  • Tableau
  • Matplotlib
Portfolio

Key Database & Software Projects Developed

Mangrove Infoline Database

A specialized web resource for mangrove plants and protein sequence information.

www.manmedinfoline.in

FERN Ethnomedicinal Plant Database

A platform organizing fern ethnomedicinal plant knowledge to support computational drug-discovery research.

www.ferndatabase.in

LegumeDB

An ethnomedicinal database and comparative evolutionary resource for legume matK proteins.

www.legumedatabase.co.in

MPRDB

Mangrove Plant Resource Database for sequence and structural evaluation.

Listed domains are project references; current online availability may vary.
Research

Peer-Reviewed Publications

  1. 2021

    Recent advances in gas-mediated cancer therapy: a review

    Thakar, S. B., Ghorpade, P. N., Shaker, B., Lee, J., & Na, D.

    Environmental Chemistry Letters, 19, 2981–2993

  2. 2020

    Full-length transcriptome sequencing provides insights into the evolution of apocarotenoid biosynthesis in Crocus sativus

    Yue, J., Wang, R., Ma, X., Liu, J., Lu, X., Thakar, S. B., An, N., Liu, J., Xia, E., & Liu, Y.

    Computational and Structural Biotechnology Journal, 18, 774–783

  3. 2020

    Bioinformatics Analysis of The Rhizosphere Microbiota of Dangshan Su Pear in Different Soil Types

    Ma, X., Thakar, S. B., Zhang, H., Yu, Z., Meng, L., & Yue, J.

    Current Bioinformatics, 15(5)

  4. 2017

    LegumeDB: Development of Legume Medicinal Plant Database and comparative molecular evolutionary analysis of matK proteins of legumes and mangroves plants

    Thakar, S. B., & Sonawane, K. D.

    Current Nutrition & Food Science

  5. 2016

    Phylogenetic, Sequence Analysis and Structural Studies of Maturase K Proteins from Mangroves

    Thakar, S., Dhanavade, M., & Sonawane, K.

    Current Chemical Biology, 10(2), 135–141

  6. 2015

    FERN Ethnomedicinal Plant Database: Exploring fern ethnomedicinal plants knowledge for computational drug discovery

    Thakar, S., Ghorpade, P., Kale, M., & Sonawane, K.

    Current Computer-Aided Drug Design, 11(3), 266–271

  7. 2013

    Mangrove Infoline Database: A database of mangrove plants and protein sequence information

    Thakar, S. B., & Sonawane, K. D.

    Current Bioinformatics, 8, 524–529

Contact

Let's discuss your research.

Share your research objective and the data you have.

  • Define the question — objective, study design, data requirements and scope.
  • Review the data — quality, permissions, missing values and potential sources of bias.
  • Explain & deliver — interpretable figures, reports and agreed reproducible materials.
Please share project details only. Do not include patient identifiers, credentials or confidential raw datasets.

Tell us about your next study.

Share your research objective and the data you have.