Hi, I’m Mehdi Fekih
I am a Data Scientist. Data Engineer. Developer. Home Automation Specialist.My Blog
How to Back Up Your Personalized AI Agent to GitHub
The more you use an AI agent, the more valuable it becomes. You customize its personality, add skills, teach it how you work, and build up memory and context. If the machine running it fails, rebuilding all of that from scratch would be painful.
A practical safety net is to back up the files that make your agent personal into a private GitHub repository — while carefully keeping secrets, tokens, and credentials out of it.
The golden rule: back up the agent’s identity, not the keys to your accounts. A private repo is a backup location, not a password vault.
What to back up
The goal is simple: (your personalized agent) − (generic install) − (secrets). Concretely, that usually means:
- Configuration files — model, provider, personality, terminal preferences, and other agent-level settings.
- Personality / instruction files — the role, tone, and rules that define how the agent behaves.
- Memory files — long-term context and preferences. Review them carefully first: they may contain personal or identifying information.
- Scheduled tasks — cron jobs and recurring automations that are easy to forget during a rebuild.
- Custom skills — scripts and workflows you created or modified.
What never belongs in the repo
.envfiles- API keys and access tokens
- OAuth and authentication files
- Private URLs, chat IDs, or other sensitive identifiers
The backup workflow
1. Create a private GitHub repository
Create a new repo dedicated to the backup and set visibility to Private. Initialize it empty — no README, no .gitignore, no license — so the agent can set it up cleanly.
2. Generate a fine-grained access token
In GitHub, go to Settings → Developer settings → Personal access tokens → Fine-grained tokens. Restrict it to the backup repository only, with Contents: Read and write permission. Set an expiration date if you can — a leaked token with an expiry does far less damage.
3. Store the token securely
Copy the token into a password manager. Never paste it into files that will be committed. If it is ever exposed, revoke it in GitHub immediately and generate a replacement.
4. Connect the agent
Add the token through your agent’s configuration so it is stored as a secret, not inside the files you plan to back up. For example:
agent config set GITHUB_TOKEN <your-token>
5. Run the first backup manually
Ask the agent to commit the configuration, memory, skills, and scheduling files. Then open the repo and review every committed file — confirm that no keys, tokens, or private notes slipped in.
6. Automate recurring backups
Once the first backup is clean, schedule a daily job and a notification after each run. Keep reviewing occasionally: automation is helpful, but it should not be blind.
Final thoughts
A personalized AI agent becomes hard to replace once it holds your settings, memory, and workflows. A clean private GitHub backup gives you a restore path when hardware fails or you migrate machines — without turning GitHub into a secrets manager.
Inspired by a guide from Zawanah about backing up a Hermes agent to GitHub.
Obsidian et ontologies pour des workflows LLM efficaces
Comment utiliser Obsidian comme base de connaissances structurée et des ontologies pour améliorer la récupération, le prompting et la cohérence des grands modèles de langage — outils, templates et exemples opérationnels.
Obsidian peut dépasser le rôle de simple carnet : bien structuré, il devient le noyau d’un workflow RAG fiable. Ce guide présente, de façon pragmatique, pourquoi une ontologie réduit les hallucinations, comment la traduire dans Obsidian et quels paramètres tester pour un mécanisme de récupération robuste. Vous trouverez des templates YAML, des requêtes Dataview, et une checklist de déploiement.

Les ontologies expliquées et pourquoi elles réduisent les hallucinations
Si vous voulez que vos workflows LLM cessent d’inventer des faits, structurez votre connaissance. Une ontologie pragmatique — entités, relations, propriétés — impose des types et des contraintes qui réduisent le bruit lexical et les sauts inferentiels du modèle.
Définition rapide
Une ontologie est un schéma déclaratif : entités (p.ex. Projet, Tâche), relations (dépend_de, responsable_de) et propriétés (statut, échéance). Elle sert de filtre pour la récupération (metadata‑filtering) et comme guide pour composer des prompts factuels.
Exemple pratique et impact mesurable
Minimal : entités = Projet, Tâche, Ressource ; relations = dépend_de, responsable_de ; propriétés = statut, priorité, échéance. Dans un pilote RAG de 200 requêtes internes, l’utilisation d’un filtrage ontologique a fait passer la précision d’extraits pertinents d’environ 62% à ~86–90% (top_k=3). Le mécanisme : on limite les candidates par type/relations, donc le LLM a moins d’hypothèses à combler.
| Structure libre (notes) | Métadonnées frontmatter | Ontologie déclarative |
|---|---|---|
| Précision : faible (≈50–70%) | Précision : moyenne (≈70–80%) | Précision : élevée (≈85–95%) |
| Réutilisabilité : limitée | Réutilisabilité : bonne (recherche + tags) | Réutilisabilité : excellente (requêtes typées, RAG) |
| Coût maintenance : faible initialement, chaos à long terme | Coût maintenance : modéré | Coût maintenance : élevé (annotation, gouvernance) |
Coûts, limites et signaux d’échec
L’effort initial porte sur l’annotation, les règles de nommage et la validation métier. Démarrez avec ~10 entités et ~20 relations, puis itérez après 4–6 semaines. Signaux d’alerte : perte de précision >10 pts, échéances contradictoires, ou multiples assertions incompatibles pour la même entité. Ne complexifiez pas l’ontologie trop tôt — restez pragmatique.

Mettre en place une ontologie dans Obsidian — guide pas à pas
Étapes pragmatiques
Commencez petit et reproductible : choisissez un dossier racine (ex. /Projects), créez un Ontology Index listant types et relations, puis standardisez un template YAML pour chaque type. Activez Templates (core) ou Templater pour insérer automatiquement le frontmatter, et Dataview pour interroger ces métadonnées. Testez avec 100–500 fiches avant d’étendre.
Template YAML frontmatter (exemple)
---
type: Projet
status: active
owner: "Claire Dupont"
date: 2026-09-01
tags: [migration, Q4]
relations:
depends_on: ["Projet Beta"]
responsible_for: ["Tâche 123", "Tâche 124"]
---
Exemple Dataview — lister tâches dépendantes
TABLE status, owner, file.link as task
FROM "Tâches"
WHERE type = "Tâche" AND contains(relations.depends_on, "Projet Alpha")
SORT date asc
Prompt RAG (template opérationnel)
System: You are a helpful assistant. Use up to top_k=5 sources (chunk_size≈500 tokens), prioritize owner and date.
Context (from retrieved sources):
- Title: {{title}} | type: {{type}} | owner: {{owner}} | date: {{date}} | excerpt: {{excerpt(200)}}
User question: {{user_query}}
Answer strictly from context; if missing, say you don't know.
Paramètres et automations recommandés
Configurez Templates/Templater pour créer rapidement des notes normalisées. Activez Dataview + MetaEdit pour manipuler le frontmatter sans ouvrir chaque note. Règles pratiques : chunk_size ≈ 500 tokens, top_k 4–8 pour support, versionnez le modèle d’embeddings et stockez la métadonnée de génération.
Checklist de déploiement (priorités)
- Définir 10 entités clés et 20 relations fréquentes (Priorité haute)
- Créer templates YAML pour chaque type (Priorité haute)
- Configurer Templates/Templater et Dataview (Priorité moyenne)
- Indexer 100–500 notes pilote et mesurer précision (Priorité moyenne)
- Former 1–2 utilisateurs métiers pour validation (Priorité haute)
- Évaluer besoin d’une vector DB externe selon volume (Priorité basse)
Intégrations et quand externaliser la vector DB
Embeddings : OpenAI pour prototypage rapide, Hugging Face pour open models. Vector DBs : Pinecone, Weaviate, Milvus en production. Règle simple : restez local si <50k vecteurs et usage mono‑utilisateur ; externalisez si vous prévoyez >100k vecteurs, usages concurrents ou besoins HA/replication.

Workflows opérationnels recommandés
RAG temps réel pour support client
Flux : ingestion → embeddings → retrieval → prompt template → post‑filtering. Paramètres concrets : chunk_size 500–1 000 tokens, upsert batch 32–128, top_k 5–8, embedding model type 1 536–3 072 dims. Latence cible retrieval <200 ms. Contremesures aux risques : versionnez les fiches (frontmatter updated_at), réindexation incrémentale 6–24 h pour contenu majeur, et human‑in‑loop sur réponses « doute ».
Pipeline batch pour synthèse de veille
Flux : scraping → nettoyage OCR/NER → classification → stockage atomique dans Obsidian → génération de résumés. Paramètres : batch scraping 500–2 000 pages, chunk_size 400–600, top_k 10 pour synthèse. Déduplication par hash + seuil de similarité (cosine ≥ 0.95) avant upsert.
Étude de cas condensée — startup SaaS (imaginaire)
Situation : 50 employés, 2 000 tickets/mois, 600 récurrents. Actions : ontologie produit (≈60 types), 2 000 fiches, embeddings OpenAI, top_k 5, post‑filtering par date/owner, rerank cross‑encoder. Résultats en 3 mois : tickets récurrents −30%, temps moyen résolution −22% (45 → 35 min), économie ≈ 110–120 heures support/mois. Coût prototype : embeddings + vectordb ≈ 800–1 200 €/mois ; ROI projeté 3–6 mois selon adoption.
Opinion : Les ontologies ne sont pas une panacée, mais elles transforment la base documentaire en référentiel interrogeable et auditables. Le véritable levier est la discipline opérationnelle : gouvernance, réindexation et indicateurs (précision RAG, temps de résolution) pour mesurer le retour.
Conclusion
Structurer l’information dans Obsidian avec une ontologie claire change la relation entre vos données et les LLM : on passe d’une saisie brute à un référentiel interrogeable et explicable. Les gains sont concrets — meilleures réponses, moins d’hallucinations, réutilisabilité — mais ils exigent modélisation et discipline technique (embeddings, chunking, plugins). Adoptez une approche itérative : commencer simple, mesurer, enrichir l’ontologie et automatiser les mises à jour.
Tarification à l’usage et économie des tokens IA, vers un nouveau paradigme pour l’intelligence artificielle API
Comment la facturation par token redéfinit l’accès, le coût et la stratégie des API d’intelligence artificielle
L’intelligence artificielle bascule vers un nouveau modèle économique : l’API à la demande, tarifiée au token. Que l’on choisisse OpenAI, Anthropic, Google, Mistral ou l’on expérimente des plateformes d’agrégation comme OpenRouter, la tarification au token applique au monde de l’IA les recettes qui ont fait le succès du cloud. Ce nouveau standard bouleverse en profondeur la façon dont développeurs, entreprises et décideurs conçoivent l’accès à l’intelligence artificielle : flexibilité accrue, paiement à l’usage, choix concurrentiels, mais aussi nouveaux risques. Décryptage des mécanismes, atouts et challenges du paradigme « API IA », où chaque token devient décision stratégique.

Du logiciel à l’API IA : naissance d’une nouvelle infrastructure
De la licence logicielle au cloud : une rupture définitive
L’histoire du numérique a été marquée par trois modèles dominants en vingt ans : la licence perpétuelle (logiciel sur site, paiement unique), l’SaaS avec abonnement, puis le cloud à la consommation. Désormais, l’intelligence artificielle ouvre une quatrième ère : celle de l’API IA facturée à l’usage, au token, qui combine la puissance du cloud et la flexibilité d’interfaces programmables.
| Modèle | Accès | Facturation | Infrastructure | Exemples |
|---|---|---|---|---|
| Licence perpétuelle | Sur site | Unique/annuelle | Serveurs internes | Microsoft Office, Photoshop |
| SaaS | En ligne | Abonnement | Cloud fournisseur | Salesforce, Slack |
| API IA par token | Programmable à la demande | À l’usage (token) | IA mutualisée cloud | OpenAI, Google Vertex AI, Mistral, Anthropic, xAI |
API IA : simplicité d’intégration, usage ajusté
L’API d’intelligence artificielle révolutionne l’accès aux LLM avancés : plus besoin de déployer ni gérer d’infrastructure lourde. Un endpoint, une clé d’authentification : en quelques minutes, l’IA enrichit une app, un chatbot, une plateforme métier.
La facturation « à la consommation » par token garantit une maîtrise immédiate des coûts : seul l’usage réel compte. Cette granularité s’accompagne de :
- Flexibilité : choix dynamique du modèle selon la tâche ou le budget.
- Implémentation rapide : ajout d’intelligence sans dettes techniques majeures.
- Mutualisation : partage des ressources sur un cloud opérant à très grande échelle.
- Évolutivité automatique : support de la montée en charge du service, sans anticipation matérielle.
« L’intelligence artificielle devient l’infrastructure invisible, comme l’électricité ou l’eau, mais à la granularité du token. »
Marché des API IA : une compétition accélérée par la standardisation
La généralisation de l’API IA par token stimule l’émergence d’un écosystème ultra-concurrentiel : OpenAI (ChatGPT, GPT-4), Anthropic (Claude 3), Google (Gemini, Vertex AI), Mistral (Fast, Large via OpenRouter), xAI (Grok)… Chacun propose désormais des APIs taillées pour la sélection à la demande, tandis que des plateformes multi-fournisseurs (OpenRouter, Together AI) permettent de combiner, comparer ou basculer de l’un à l’autre sans friction technique.
Pour les entreprises, cette standardisation réduit le verrouillage technique et accélère l’expérimentation. Toutefois, l’agilité offerte par la logique du token implique une vigilance accrue : suivi des coûts en temps réel, gestion des quotas, dépendance au réseau et à la qualité des services sous-jacents — autant de nouveaux leviers stratégiques.

Tarification au token : révolution des coûts et des usages
Unité de facturation : précision, prédictibilité et passage à l’échelle
La facturation à l’usage par token inscrit l’IA dans la lignée du « pay as you go » du cloud, avec une granularité extrême. Un token (environ 0,75 mot) devient la brique monétaire minimale de chaque appel API. Cela rend possible une mesure précise des coûts, adaptée du particulier à la multinationale.
- Paiement instantané et sans engagement : chaque prompt, chaque réponse, chaque pipeline a un coût lisible.
- Montée en charge fluide : du proof of concept à l’application grand public, la structure de coût suit l’usage effectif.
- Pression concurrentielle à la baisse : la compétition entre fournisseurs entraîne régulièrement une réduction du prix par token et encourage l’apparition d’alternatives (API open source, modèles auto-hébergés, etc.).
Comparatif : tarifs publics 2024 des principaux acteurs
| Fournisseur / Modèle | Entrée (1k tokens) |
Sortie (1k tokens) |
|---|---|---|
| OpenAI GPT-4o | 0,005 $ | 0,015 $ |
| Mistral Large | 0,008 $ | 0,024 $ |
| Anthropic Claude 3 Sonnet | 0,003 $ | 0,015 $ |
| Google Gemini 1.5 Pro | 0,0025 $ | 0,0025 $ |
| OpenRouter/agrégateurs | Selon modèle sous-jacent | |
Exemple concret : 10 000 requêtes/an avec 3 000 tokens par interaction coûtent ~900 $ sur GPT-4o, mais pourraient descendre à 120 $ avec un modèle interne Mistral Small.
Checklist : optimiser ses usages IA API au token
- Suivre en temps réel la consommation par type de prompt ou projet.
- Définir des plafonds de coûts mensuels ou par client.
- Rédiger des prompts compacts pour éviter l’explosion tarifaire liée à la verbosité.
- Comparer régulièrement les fournisseurs pour profiter de la guerre des prix.
- Automatiser l’alerting sur les dérives de consommation.
Avantages majeurs pour tous les profils métiers
- Innovation accessible : barrieres d’entrée basses, prototypage instantané.
- Évolutivité native : même modèle économique pour 10 ou 100 000 utilisateurs.
- Pas de gestion matérielle : ni serveurs, ni GPUs à anticiper.
- Mobilité contractuelle : migration ou diversification via APIs compatibles.
- Alignement immédiat coût/usage : dépenses proportionnelles à la valeur générée.
Limites et défis concrets
- Difficulté d’anticipation sur les usages massifs ou chaotiques.
- Nécessite une optimisation fine des prompts pour contenir la facture.
- Risque de « spike » (pics de charges), quotas imprévus, et interruptions de service si limites atteintes.

Risques et limites de l’économie des tokens IA
Dépendance aux fournisseurs et verrouillage technologique
L’économie du token propulse les Goliaths technologiques (OpenAI, Google, Anthropic, xAI…) au centre du jeu. Si l’API promet l’agilité, elle peut induire un lock-in propriétaire : migration coûteuse, arrêt soudain d’un modèle ou hausse imprévisible des prix sont des risques réels. Hausse tarifaire brutale de GPT-4 en 2023 : de nombreuses startups ont été contraintes d’adapter leur stratégie ou de choisir des alternatives parfois moins performantes.
Volatilité des coûts et gouvernance des budgets
À l’usage massif, la facture API IA peut s’envoler et révéler une grande instabilité financière. Des changements tarifaires soudains ou une demande fluctuante rendent la planification difficile. Comme le souligne un CTO :
« Au-delà du million de tokens par jour, notre facture API varie parfois du simple au double selon la semaine. »
Enjeux de confidentialité et concentration du marché
Envoyer des données à des API externes soulève des doutes sur la confidentialité et la traçabilité. Pour les secteurs régulés (droit, santé…), des choix alternatifs émergent : modèles open source comme Mistral ou Llama, auto-hébergement, solutions de chiffrement avancé.
Tableau de synthèse : cartographie des principaux risques
| Risque | Mécanisme | Alternatives/Pistes |
|---|---|---|
| Dépendance fournisseur | API fermées, modèles exclusifs | Open source : Mistral, Llama |
| Volatilité prix | Tarifs variables, quotas changeants | Plafonds, contractualisation long-terme |
| Confidentialité | Transit de données externes | Modèles auto-hébergés, chiffrement |
| Concentration du marché | Peu d’acteurs dominants | Plateformes multi-providers (OpenRouter) |
Face à ces risques, de nombreux professionnels privilégient une stratégie hybride : routage dynamique multi-fournisseurs et recours à des modèles locaux ou open source pour sécuriser la chaîne de valeur.

Agents intelligents et avenir API : entre diversité et concentration
L’ère des agents IA multi-modèles
L’automatisation et l’orchestration intelligente de l’accès aux API IA posent les bases d’un nouveau meta-marché. Des agents IA désormais capables de router dynamiquement les requêtes vers différents modèles selon le coût, la qualité ou les exigences réglementaires. Des plateformes comme OpenRouter ou PromptLayer offrent un accès unifié aux moteurs d’OpenAI, Anthropic, Mistral ou Google Gemini, et permettent d’associer souplesse, contrôle budgétaire et gain de temps.
Checklist stratégique : tirer profit de la diversité IA API
- Flexibilité : router en temps réel sur le modèle le plus pertinent (performance, prix, conformité).
- Suivi centralisé des coûts : analytiques intégrées, plafonds, alertes de consommation.
- A/B testing instantané : expérimentation simplifiée sur les modèles concurrents.
- Réduction du lock-in technique : migration et agrégation de fournisseurs en un clic.
« L’agent IA devient l’équivalent du load balancer pour l’IA : il module dynamiquement les ressources, libérant les développeurs de la rigidité des choix d’infrastructure. »
De la diversité à la concentration : nouveau verrouillage ?
Si l’aggrégation multi-modèles vise à réduire la dépendance, elle concentre le pouvoir entre les mains de quelques plateformes d’orchestration. Ces intermédiaires-clés, capables à terme d’imposer leurs propres règles ou surcoûts, peuvent constituer un pivot stratégique autant qu’une nouvelle source de verrouillage.
| Avantages | Risques |
|---|---|
| Indépendance vis-à-vis d’un seul fournisseur | Capture possible du marché par quelques orchestrateurs |
| Adoption rapide des innovations IA | Imposition de règles tarifaires ou techniques |
| Comparaison transparente du coût par token | Concentration accrue sur l’accès centralisé |
La diversité offerte par ces solutions est un moteur d’innovation, mais appelle à une veille stratégique permanente : garantir la souveraineté sur ses flux de données, maintenir la maîtrise budgétaire, anticiper les effets de concentration du marché.
Conclusion
La tarification à l’usage par token redéfinit l’économie de l’IA, remplaçant la licence ou l’abonnement par le calcul à l’interaction. Ce nouveau paradigme, tout en apportant agilité, mutualisation et ouverture, exige rigueur et stratégie dans la gestion des coûts, le choix des partenaires et la maîtrise de la confidentialité. L’économie des tokens façonnera la décennie à venir, autant source d’opportunités que de nouveaux équilibres à inventer entre diversité, indépendance et efficacité.
Awesome Public Datasets: A Treasure Map for Data-Driven Projects
If you have ever started a machine learning project, built a dashboard, written a research paper, or simply wanted to explore real-world data, you already know the hardest part is often not the code. It is finding a good dataset.
That is why the GitHub repository Awesome Public Datasets is such a valuable resource. It is a curated collection of public datasets organized by topic, making it easier for developers, researchers, analysts, students, and data enthusiasts to discover high-quality data sources without spending hours searching across the web.
Whether you are looking for climate records, economic indicators, social network graphs, image datasets, government data, healthcare resources, or machine learning benchmarks, this repository acts like a map to the public data ecosystem.
What Is Awesome Public Datasets?
Awesome Public Datasets is an “awesome list” dedicated to topic-centric public data sources. Like other awesome lists on GitHub, its goal is simple: collect useful links in one place and organize them so people can find what they need quickly.
The repository includes datasets from a wide range of domains, including agriculture, biology, chemistry, climate and weather, cybersecurity, economics, education, energy, finance, GIS and geospatial data, government, healthcare, image processing, machine learning, natural language processing, neuroscience, physics, social sciences, software, sports, time series, and transportation.
It also includes complementary collections, which can lead users to even more dataset repositories and archives. In short, it is not a single dataset. It is a gateway to hundreds of datasets.
Why This Repository Is So Useful
The internet is full of data, but not all data is easy to find, clean, documented, or usable. Many valuable datasets are buried inside university pages, government portals, academic archives, old project websites, or research labs.
Awesome Public Datasets helps solve that discovery problem. Instead of searching Google for “free public dataset for network analysis” or “open agriculture data,” you can browse a categorized list and quickly find relevant sources.
For example, the repository points to well-known resources such as the Stanford Large Network Dataset Collection for graph and network research, Open Food Facts for food product data, NBER Patent Citations for economics and innovation research, DIMACS Road Networks Collection for transportation and graph algorithms, and many climate, biology, finance, and machine learning datasets.
Each listing usually includes a short description and a link to the original data source. Many entries also include metadata links from the repository’s companion project, apd-core.
A Dataset Directory for Many Audiences

For Data Scientists
Data scientists can use it to find datasets for exploratory analysis, predictive modeling, visualization, and portfolio projects. Instead of working with the same few beginner datasets repeatedly, they can explore more specialized real-world data.
For Machine Learning Engineers
Machine learning engineers can find benchmarks and domain-specific data for experimenting with models. Categories like image processing, natural language, time series, and cybersecurity are especially useful for ML workflows.
For Researchers
Researchers can discover public data sources related to biology, physics, neuroscience, social sciences, climate, economics, and more. The repository can serve as a starting point for literature reviews, reproducible experiments, or interdisciplinary research.
For Students
Students learning Python, R, SQL, data visualization, or statistics can use the repo to find project ideas. Real datasets make learning more meaningful because they contain imperfections, surprises, and domain context.
For Journalists and Analysts
Data journalists and analysts can use the collection to locate public-interest datasets, especially in government, economics, transportation, education, healthcare, and climate.
What Makes It Better Than a Random List of Links?
The strength of Awesome Public Datasets is not just that it contains many links. It is the organization and curation.
The datasets are grouped by domain, so browsing feels natural. If you are interested in geospatial analysis, you can jump to GIS. If you are researching transportation networks, you can explore Transportation or Complex Networks. If you are looking for NLP resources, there is a Natural Language section.
The repository also uses status icons for entries. Some links are marked as healthy, while others are marked as needing attention. That is important because public dataset links often break over time. Seeing that maintenance status gives users a quick signal about whether a resource may need verification.
Another important detail: the README notes that the repository is automatically generated by apd-core. Contributors are asked not to edit the generated README directly, but to contribute through the appropriate metadata workflow. That makes the project more structured than a hand-edited list.
Great Project Ideas Using Awesome Public Datasets
- Build a climate dashboard: Use climate and weather datasets to visualize temperature changes, rainfall patterns, or extreme weather events over time.
- Analyze transportation networks: Explore road network datasets or public transport data to study shortest paths, congestion, or urban accessibility.
- Create a food product explorer: Use Open Food Facts or other agriculture and food datasets to analyze nutrition, ingredients, product origins, or labeling trends.
- Study social networks: Use graph datasets from the social networks or complex networks sections to learn network analysis, centrality, community detection, and graph visualization.
- Practice time series forecasting: Find time series datasets related to energy, finance, climate, or transportation and build forecasting models.
- Train an image classification model: Browse image processing datasets and experiment with computer vision techniques.
- Investigate public policy questions: Government, education, economics, healthcare, and social science datasets can support projects around inequality, public spending, population trends, or policy outcomes.
If you want help turning one of these datasets into a production-ready pipeline — ETL, storage, or a deployed model — check out my data engineering services or get in touch.
A Few Things to Keep in Mind

Awesome Public Datasets is a directory, not a guarantee that every dataset is ready to use immediately.
Before starting a project, you should always check the dataset license, whether the data is free or requires payment, update frequency, file format, documentation quality, privacy or ethical considerations, whether the link is still active, and whether the dataset is suitable for commercial use.
The repository itself notes that most datasets are free, but some are not. That distinction matters, especially for production or commercial projects.
Also, because many datasets come from third-party sources, quality can vary. Some may be clean and well-documented, while others may require significant preprocessing.
Why Public Datasets Matter
Public datasets are one of the foundations of modern data work. They make research more transparent, help students learn by doing, allow developers to test ideas, and enable journalists and citizens to investigate important questions.
Open data also lowers the barrier to innovation. A student with a laptop can analyze climate trends. A developer can build a prototype with public transportation data. A researcher can compare results using shared benchmarks. A startup can validate an idea before collecting proprietary data.
Repositories like Awesome Public Datasets make that ecosystem easier to navigate.
Final Thoughts
Awesome Public Datasets is one of those GitHub repositories worth bookmarking immediately. It saves time, sparks ideas, and opens doors to data from dozens of fields.
If you are learning data science, building machine learning models, writing research, creating visualizations, or looking for your next portfolio project, this repository is an excellent place to start.
The next time you ask, “Where can I find a good dataset?”, start with Awesome Public Datasets. And if you need a hand turning that data into something real — a dashboard, a model, a pipeline — I do this for a living. Let’s talk.
Bitcoin and “Cyclicality”: Reassuring Myth or Serious Analysis?
For several years, much of the discussion around Bitcoin has rested on one central idea: cyclicality.
Halvings, four-year cycles, an “inevitable” bull run, and an equally expected bear market. The narrative is well-oiled, almost comforting.
But one question deserves to be asked plainly:
Can we really call it serious analysis when we are waiting for a phenomenon that is supposedly obvious and predictable?
What supporters of cyclicality argue
Defenders of this view mainly rely on three elements:
- Historical data: since 2012, the major bullish phases have followed halvings.
- Programmed scarcity: the reduction in issuance is supposed to mechanically influence price.
- The repetition of human behavior: euphoria, excess, correction, forgetting, then return.
Taken individually, these elements are not absurd. The problem begins when they are presented as an almost deterministic mechanism.

Where the reasoning becomes fragile
1. A ridiculously small statistical sample
Bitcoin has existed for just over fifteen years.
Speaking of robust cycles based on three or four occurrences is closer to storytelling than science.
In finance, no one would describe that as a usable time series with a high level of confidence.
2. Confusing correlation with causation
The fact that rallies followed halvings does not prove that:
- the halving is their main cause,
- or that the same pattern will repeat identically.
The markets of 2013, 2017, and 2021 had nothing in common in terms of liquidity, participants, regulation, or macroeconomics.
3. A self-fulfilling prophecy
The more an idea is repeated, the more it influences behavior.
- Investors buy “before the halving”
- The media amplify the narrative
- Flows become synchronized
The cycle then becomes a social artifact, not a market law.
It works… until the day it no longer does.
Technical analysis: tool or illusion of control?
Technical analysis is not useless in itself. It is effective for:
- reading collective behavior,
- identifying liquidity zones,
- managing short- or medium-term risk.
But it does not turn an asset as young, political, and narrative-driven as Bitcoin into a predictable metronome.
Believing otherwise means confusing reading the past with the ability to forecast.

What Bitcoin really is
Bitcoin is not:
- a stock with cash flows,
- a traditional commodity,
- a mature asset.
It is all at once:
- a technological object,
- an experimental monetary asset,
- an ideological symbol,
- a field for global speculation.
Reducing all of that to a simple cyclical curve is intellectually comfortable, but analytically poor.
So, is skepticism a mistake?
No.
Being skeptical of cyclicality presented as obvious is, on the contrary, a way to:
- reject overly neat narratives,
- avoid lazy certainties,
- maintain an open analytical stance.
The real danger is not doubting cycles.
The real danger is mistaking them for natural laws.
Conclusion
Bitcoin cyclicality is a useful narrative, sometimes effective, but never guaranteed.
It helps structure expectations, not predict the future.
In a market this young and shifting, the only truly rational position remains intellectual caution.
The day the “obvious” cycle fails, it will not be an anomaly.
It will simply be the market reminding us that it owes nothing to our charts.
Configuring Amazon S3 Access Keys Securely for UpdraftPlus (WordPress)
Using Amazon Web Services S3 as remote storage for UpdraftPlus is a common and reliable approach for backing up a WordPress site.
However, many users are confused when AWS warns against long-term access keys and suggests alternative authentication methods.
This article explains the correct and secure way to configure S3 access for UpdraftPlus, why AWS shows those warnings, and what best practices actually apply in a real WordPress environment.
Why AWS Warns About Long-Term Access Keys
AWS strongly encourages modern authentication mechanisms such as:
- IAM Roles
- Temporary credentials (STS)
- Workload identity federation
These are excellent practices when your application runs inside AWS (EC2, ECS, Lambda).
However, classic WordPress hosting does not support IAM roles.
If your WordPress site runs on:
- Shared hosting
- A VPS (DigitalOcean, OVH, Hetzner, etc.)
- On-premise infrastructure
then access keys are the only supported and correct solution.
AWS warnings are contextual, not prohibitions.
The Correct AWS Use Case for UpdraftPlus
When creating an access key in IAM, AWS asks you to select a use case.
✅ Correct choice:
Application running outside AWS
This matches the reality:
- UpdraftPlus is a third-party PHP application
- It runs outside AWS
- It only needs programmatic access to S3
Selecting this option does not weaken security and does not change how credentials work.
It simply helps AWS categorize usage internally.
Secure Architecture Overview
WordPress
└─ UpdraftPlus
└─ IAM User (restricted)
└─ S3 Bucket (private)
Key principle: least privilege.
Step 1 – Create a Dedicated S3 Bucket
Best practices:
- Private bucket (Block all public access)
- Dedicated to backups only
- Optional versioning enabled
- Optional lifecycle rules (auto-delete old backups)
Example bucket name:
my-wp-backups-prod
Step 2 – Create a Dedicated IAM User
Never use:
- Root credentials
- A shared IAM user
- Broad policies like
AmazonS3FullAccess
Create a single-purpose IAM user, for example:
updraftplus-wordpress
Programmatic access only.
Step 3 – Attach a Minimal IAM Policy
This policy allows only what UpdraftPlus needs:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "UpdraftPlusS3Access",
"Effect": "Allow",
"Action": [
"s3:PutObject",
"s3:GetObject",
"s3:DeleteObject",
"s3:ListBucket"
],
"Resource": [
"arn:aws:s3:::my-wp-backups-prod",
"arn:aws:s3:::my-wp-backups-prod/*"
]
}
]
}
This prevents:
- Access to other buckets
- Account-wide damage if credentials leak
Step 4 – Generate Access Keys
Generate:
- Access Key ID
- Secret Access Key
Store them securely and never commit them to Git.
Add a description such as: UpdraftPlus – production backups
Step 5 – Configure UpdraftPlus
In WordPress:
- Settings → UpdraftPlus → Settings
- Select Amazon S3
- Enter:
- Access Key ID
- Secret Access Key
- Bucket name
- Region
- Set a bucket subpath (recommended):
wordpress/site-prod/
- Save and test the connection
If it fails, the cause is almost always:
- Wrong region
- Incorrect IAM policy
- Typo in bucket name
Optional Hardening (Recommended)
Lifecycle Rules
Automatically delete old backups:
- Daily: keep 14 days
- Monthly: keep 6 months
This avoids silent storage cost growth.
Server-Side Encryption
Enable default SSE-S3 (AES-256).
No code changes required.
IP Restriction (Advanced)
If your hosting provider has a static IP, you can restrict IAM access to that IP range.
Common Mistakes to Avoid
- Using root access keys
- Granting
AmazonS3FullAccess - Making the bucket public
- Skipping lifecycle rules
- Reusing credentials across multiple sites
Final Verdict
For WordPress + UpdraftPlus:
- Long-term access keys are normal
- AWS warnings are generic
- Least-privilege IAM policies are what actually matter
Used correctly, this setup is secure, stable, and industry-standard.
What I Can Do For You
Data Science
Unlocking insights and driving business growth through data analysis and visualization as a freelance data scientist.
Data Analysis
Helping businesses make informed decisions through insightful data analysis as a freelance data analyst.
Website Development
Bringing your online presence to life with customized website development solutions as a freelance developer.
Home Automation
Transforming your living space into a smart home with custom home automation solutions as a freelance home automation expert.
Docker & Server
Optimizing your software development and deployment with Docker and server management as a freelance expert
Consultancy
Identification of scope, assessment of feasibility, cleaning and preparation of data, selection of tools and algorithms.
My Portfolio
My Resume
Education
Msc Data Science And Artificial Intelligence
2022 - 2023Training in data science & artificial intelligence methods, emphasizing mathematical and computer science perspectives.
Master In Management
EDHEC Business School (2005 - 2009)English Track Program - Major in Entrepreneurship.
BSc in Applied Mathematics and Social Sciences
University Paris 7 Denis Diderot (2006)General university studies with a focus on applied mathematics and social sciences.
Education
Higher School Preparatory Classes
Lycée Jacques Decour - Paris (2002 - 2004)Classe préparatoire aux Grandes Écoles de Commerce. Science path.
Scientific Baccalaureate
1999 - 2002Mathematics Major
Data Science
Python
SQL
Machine learning libraries
Data visualization tools
DESIGBig Data (Spark, Hive)
Data Analysis
Spreadsheet software
Data visualization tools (Tableau, PowerBI, and Matplotlib)
Statistical software (SAS, SPSS)
SAP Business Objects
Database management systems (SQL, MySQL)
Development
HTML
CSS
JAVASCRIPT
SOFTWARE
Version Control Systems
MLOps
CI/CD
Docker and Kubernetes
AutoML
Model serving frameworks
Prometheus, Grafana
Job Experience
Consulting, Automation and Security
(2019 - Present)Implementation of automated reporting tools via SAP Business Objects, Processing and securing sensitive data (data wrangling, encryption, redundancy), Remote monitoring management solutions via connected objects (IoT) and image processing, Internal pentesting and network security consulting, VPN implementation, Outsourcing of servers
Consulting, E-commerce and Digital Marketing
(2017 - 2019)Consulting in e-commerce and digital marketing (Bangkok area), Booking.com, Airbnb, Agoda online booking management for third parties, SEO in the hotel industry.
Consulting, internal company network
(2016 - 2017)Implementation of corporate networks and virtualization solutions (rack cabling, firewalls, proxmox virtualization), Management of firewalls and internal networks. (pfSense)
Entrepreneurship Experience
Founder, web developer
(2015 - 2020)Programming and maintenance of websites and web applications, Consulting in digitalization and process optimization for local SMEs, Implementation of turnkey e-commerce solutions.
Corporate Banking
Assistant Fund Manager
Credit Portfolio Management - CALYON - 2008● Preparation of committee notes for new ABS/CDO credit derivative investments ● Calculation and measurement of portfolio risk (Value-at-Risk, exotic and vanilla ABS, SWAP, liquidity lines) ● Daily monitoring of credit derivatives portfolio structures (Mark-to-Market, P&L, re-financing) ● Design of risk measurement and decision support tools in VBA.
Credit Risk Analyst
Risk and Controls Department – NATIXIS – ParisStudy of financing files for review by the credit committee (structured finance, commodity trade finance, and car manufacturers) ● Financial analysis, rating, and credit risk analysis of a portfolio of companies ● Financing files studied: from €1m to €1000m
Retail Banking
Assistant Business Account Manager
BNP Paribas - 2006● Writing reports on business plans for small SMEs ● Risk and feasibility Analysis and Decision making ● Negotiation of financing xpackages with applicants
Pierre Toul
Chief Technical OfficerData & Process Automation
Jan. 2025 – Sep. 2025Mehdi a su rapidement comprendre nos processus métiers et transformer des outils Excel/VBA existants en solutions Python plus robustes et maintenables. Son autonomie, sa rigueur et son attention portée aux utilisateurs ont facilité l'intégration et l'adoption des nouveaux outils.
Arkadus Romitry
Engineering Data LeadData Engineering & Analytics
Sep. 2024 – Jan. 2025Mehdi a mené une analyse complexe sur plusieurs années de données techniques et a su en extraire des tendances utiles aux équipes métier. Il combine efficacement maîtrise technique, rigueur analytique et capacité à restituer des résultats complexes de manière claire.
Contact Me
Mehdi Fekih
Data Scientist.I am available for freelance work. Connect with me via this contact form or feel free to send me an email.
Phone: +33 (0) 7 82 90 60 71 Email: mehdi.fekih@edhec.com