Skip to content
View huseyincenik's full-sized avatar
🎯
Focusing
🎯
Focusing

Highlights

  • Pro

Block or report huseyincenik

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
huseyincenik/README.md

πŸ’« About Me

Data Scientist with a strong Mathematics & Statistics foundation and 10+ years in education. I turn complex, unstructured data into decisions β€” NLP & LLM systems, RAG pipelines, large-scale web scraping & automation, and dashboards teams actually use. Let's analyze the data together πŸš€

Focus areas β€” end-to-end NLP & LLM systems Β· RAG Β· fine-tuning Β· evaluation Β· large-scale web scraping & automation Β· agentic workflows in production Β· dashboards that drive decisions Β· πŸ† 2024 Tableau Ambassador

🌐 Socials

LinkedIn Kaggle Tableau Hugging Face Streamlit Medium ResearchGate ORCID


πŸ› οΈ Technical Ecosystem

Languages & Query
Python β€” general-purpose language for data science & automationSQL β€” querying & modelling relational dataR β€” statistical computing & analysisJavaScript β€” dashboards, tooling & web integrationsBash β€” shell scripting & automation on Linux serversHTML5 β€” report & dashboard markupCSS β€” styling reports & Streamlit / web apps
Python (Pandas Β· NumPy Β· SciPy) Β· SQL Β· R Β· JavaScript Β· Bash Β· HTML5 Β· CSS Β· Regex
ML & Deep Learning
scikit-learn β€” classical ML: classification, regression, clusteringTensorFlow β€” deep-learning models & trainingKeras β€” high-level neural-network APIPyTorch β€” deep learning & fine-tuningMLflow β€” experiment tracking & model lifecyclepandas β€” dataframes, cleaning & feature engineeringNumPy β€” numerical arrays & linear algebraSciPy β€” scientific computing & statistics
scikit-learn Β· TensorFlow Β· Keras Β· PyTorch Β· MLflow Β· Pandas Β· NumPy Β· SciPy β€” classification, regression, clustering, model evaluation & tuning
LLM & NLP
LangChain β€” building LLM & RAG pipelines and agentsHugging Face β€” transformer models, datasets & fine-tuningAnthropic (Claude) β€” LLM API for production workflowsOpenAI β€” GPT models & embeddings APIspaCy β€” industrial NLP: NER, tokenisation, pipelinesLLMs β€” prompt engineering, evaluation & guardrailsRAG β€” retrieval-augmented generation over vector stores
LangChain Β· Hugging Face Β· Anthropic Β· OpenAI Β· spaCy / NER Β· LLMs Β· RAG Β· agents, prompt engineering & evaluation frameworks
Data Engineering
Scrapy β€” large-scale web-scraping frameworkSelenium β€” browser automation for dynamic sitesBeautifulSoup β€” HTML parsing & extractionSupabase β€” Postgres backend, auth & storagePostgreSQL β€” primary relational databaseMongoDB β€” document database for semi-structured dataMicrosoft SQL Server β€” enterprise relational databaseElasticsearch β€” full-text search & analytics engineTypesense β€” fast, typo-tolerant search engineApache Kafka β€” event streaming & data ingestion
Scrapy Β· Selenium Β· BeautifulSoup Β· Supabase Β· PostgreSQL Β· MongoDB Β· MS SQL Server Β· Elasticsearch Β· Typesense Β· Apache Kafka β€” automated pipelines & REST APIs
BI & Visualization
Tableau β€” interactive dashboards (2024 Tableau Ambassador)Power BI β€” business dashboards & data modellingLooker Studio β€” free cloud dashboards & reportsMatplotlib β€” programmatic plotting & figuresPlotly β€” interactive chartsGoogle Sheets β€” lightweight analysis & sharingExcel β€” spreadsheets, pivots & ad-hoc analysis
Tableau Β· Power BI Β· Looker Studio Β· Matplotlib Β· Plotly Β· Google Sheets Β· Excel β€” plus Seaborn
Cloud, Deploy & Apps
AWS β€” Lambda, S3, EC2, SageMaker deploymentsGoogle Cloud PlatformDocker β€” containerised services & reproducible envsLinux β€” server administration (Hetzner, Docker hosts)FastAPI β€” async APIs for ML & LLM servicesStreamlit β€” data apps & model demosGradio β€” quick ML model UIsDokploy β€” self-hosted PaaS for deploying apps & databasesReact Native β€” cross-platform mobile apps
AWS (Lambda Β· S3 Β· EC2) Β· GCP Β· Docker Β· Linux Β· FastAPI Β· Streamlit Β· Gradio Β· Dokploy Β· React Native

πŸ“ˆ Professional Highlights

  • Data Science: Build agentic workflows (Tagger, Extractor, Consolidator) and fine-tune models for niche sectors like energy and clinical data.
  • Automation: Ship end-to-end automation systems β€” large-scale data extraction, trade-data pipelines and automated content management via REST APIs.
  • Recognition: 2024 Tableau Ambassador; contributed to The Globby's πŸ† 2nd place at the WTO Small Business Champions 2026.

πŸ—“οΈ Career Timeline

Math Teacher β€” Private lessons (Sep 2012 – Present)Upwork β€” Data Scientist / Data Analyst (Mar 2023 – Present)EnkiAI β€” Data Scientist β†’ Senior Data Scientist (Feb 2024 – Present)The Globby β€” Senior Data Scientist (Feb 2025 – Present)John Snow Labs β€” Data Scientist (Dec 2025 – Present)Westaco β€” Senior Data Scientist (Feb 2026 – Apr 2026)AI PR Reviewer β€” Creator (Jul 2026 – Present)

Open ↴  Math Teacher Β· Upwork Β· EnkiAI Β· The Globby Β· John Snow Labs Β· Westaco Β· AI PR Reviewer


πŸ’Ό Professional Experience

  • Creator | AI PR Reviewer Β· GitHub App (Jul 2026 – Present)

    • Designed and shipped a free, bring-your-own-key GitHub App that automatically reviews every pull request for bugs, security issues, and performance problems β€” running entirely on GitHub Actions and costing $0 to operate, since it calls OpenRouter's free LLM models with each user's own key instead of a metered backend. Built the full pipeline solo: GitHub App auth, a Cloudflare Workers setup dashboard with a live free-model picker, encrypted secret provisioning, and automatic model fallback for rate-limited requests. Install it β†’
  • Senior Data Scientist | Westaco Β· Freelance Β· Romania Β· Remote (Feb 2026 – Apr 2026)

    • AI-powered chatbot development and LLM integrations. Optimized the end-to-end chatbot architecture (+15% response accuracy, βˆ’20% latency), integrated multi-provider LLM APIs (OpenAI, Anthropic) with fallback mechanisms (99.5% uptime), built async services with FastAPI / Flask, and implemented RAG pipelines on ChromaDB / Pinecone (+22% factual retrieval accuracy) with structured output guardrails.
  • Data Scientist | John Snow Labs Β· Contract Β· Remote (Dec 2025 – Present)

    • Led end-to-end data curation for healthcare and clinical text datasets (+25% pipeline efficiency); cleaned, normalized and annotated large-scale medical text (+30% consistency); labeled clinical entities with domain-specific frameworks (+20% NER, classification and information-extraction accuracy); ran QA reviews, validation checks and feedback loops (+25% annotation accuracy and compliance) for NLP and LLM model development.
  • Senior Data Scientist | The Globby Β· Full-time Β· Remote (Feb 2025 – Present)

    • Architected scalable web-scraping bots and automated pipelines in Python (βˆ’75% manual reporting time); integrated UN Comtrade and Trademap APIs into Supabase across 240+ countries; engineered LLM-driven B2B buyer targeting and automated smart email generation β€” contributing to the platform's πŸ† 2nd place at the WTO Small Business Champions 2026. Managed Hetzner (Linux/Docker) infrastructure at 99.9% uptime.
  • Data Scientist β†’ Senior Data Scientist | EnkiAI Β· Full-time Β· Remote (Feb 2024 – Present)

    • Own NLP and LLM-based systems end-to-end for the energy sector, from experimentation to production. Design, optimize and scale Retrieval-Augmented Generation (RAG) pipelines (+20% response relevance), define evaluation and monitoring frameworks for fine-tuned models, lead data-labeling and QA, and mentor junior data scientists. Promoted to Senior Data Scientist in Sep 2025.
  • Data Scientist / Data Analyst | Upwork Β· Freelance Β· Remote (Mar 2023 – Present)

    • End-to-end ML for clients β€” Customer Segmentation, Churn Prediction, Sentiment Analysis, RFM and Cohort analysis β€” using Logistic / Linear Regression, Random Forest, Gradient Boosting, clustering and deep learning (TensorFlow, Keras). Model deployment via Streamlit, AWS SageMaker / EC2 / Lambda / boto3; large-scale web scraping (Scrapy, Selenium, BeautifulSoup); Power BI dashboards; 1M+ records handled with Supabase and external APIs.
  • Mathematics Teacher | Private Lessons Β· Part-time Β· Erzurum / Giresun Β· Hybrid (Sep 2012 – Present)

    • 14 years designing personalized, data-informed lesson plans from student-performance analysis β€” +40% overall success, +28% course success, and a βˆ’25% mathematics failure rate.

πŸŽ“ Education

AtatΓΌrk University β€” B.Sc. in Mathematics Education (2012–2017)Mathematics teaching & pedagogical leadership (since 2012)Giresun University β€” M.Sc. in Statistics (2023–2026)

  • πŸŽ“ M.Sc. in Statistics β€” Giresun University (2023 – 2026)
    Thesis β€” Modeling of Data Obtained by Web Scraping Using ML Methods and Performance Comparison: benchmarked 16 models (Transformers, RNNs, classical ML) on a 20K-article web-scraped dataset; BART scored best (F1 0.7955). Β Repo β†’
  • πŸŽ“ B.Sc. in Mathematics Education β€” AtatΓΌrk University, Erzurum (2012 – 2017)
  • 🏫 14 years teaching Mathematics & pedagogical leadership (since 2012)

πŸ“‚ Projects

Β Data Analysis Β· Python
Cleaning, EDA, visualisation and statistical analysis in Python.
Β Data Analysis Β· SQL
Extracting insights and trends straight from SQL queries.
Β Power BI
Data modelling, interactive dashboards and report challenges.
Β Tableau
Interactive dashboards and visual data exploration.
Β Looker Studio
Dynamic cloud dashboards and customisable reports.
Β Spreadsheets
Data manipulation and analysis with spreadsheet tooling.
Β NLP
Sentiment analysis, text classification and language modelling.
Β Machine Learning
Classification, regression and clustering implementations.
Β Deep Learning
Neural networks for image recognition and NLP tasks.
Β Generative AI
Text generation, image synthesis and applied GenAI experiments.
Β React Native
Cross-platform mobile apps: components, navigation, state.

πŸš€ Live Apps

Β NER with GLiNER
Extract any entity from unstructured text β€” a GLiNER-powered NER app.
β–Ά Live app Β Β·Β  Code
Β Chat with your PDFs
Conversational Q&A over uploaded PDFs with Gemini + OpenAI via LangChain.
β–Ά Live app Β Β·Β  Code
Β Auto Analytics
Car-price estimation with Linear / Ridge / Lasso, Random Forest and XGBoost.
β–Ά Live app Β Β·Β  Code

πŸ“Š GitHub Stats

GitHub profile summary

Top languages by repo Β  Top languages by commit

GitHub stats Β  Most productive time

πŸ“Œ Kaggle

Kaggle Competitions: Contributor (2 entered)Kaggle Datasets: Expert (rank 1,568 / 11k)Kaggle Notebooks: Expert (rank 1,496 / 61k)Kaggle Discussion: Contributor (80 posts)

full Kaggle profile β†’


@huseyincenik Β· LinkedIn Β· Tableau Β· Kaggle
Profile views

Pinned Loading

  1. multi_agent_rag_system multi_agent_rag_system Public

    A self-improving multi-agent retrieval-augmented generation (RAG) API. A query enters, several specialised agents collaborate, and a final answer is streamed back β€” with full execution traces, stru…

    Python 3

  2. nlp_natural_language_processing nlp_natural_language_processing Public

    NLP (Natural Language Processing)

    Jupyter Notebook 5 2

  3. markdown_html_reference_catalog markdown_html_reference_catalog Public

    Markdown Html Reference Catalog

    Jupyter Notebook