batchCooking/apps/api/src/sources/manger-bouger.ts
kyuno053 5ea1026151
feat(recipes): adaptateurs RecipeSourceAdapter pour Marmiton, 750g et Manger Bouger (#67)
* feat(recipes): ajoute un adaptateur RecipeSourceAdapter pour Marmiton

Étend jsonLdRecipeAdapter (json-ld-recipe.ts) plutôt que de dupliquer sa
logique : marmitonAdapter délègue fetchDetail/parse directement à
l'adaptateur générique JSON-LD (une page recette marmiton.org expose un
Recipe schema.org standard), et n'ajoute que ce que l'adaptateur
générique ne peut pas offrir — un list() qui lit l'ItemList schema.org
embarqué sur la page de résultats de recherche de marmiton.org (pagination
via &page=N, fin de résultats détectée via la réponse 404 renvoyée
au-delà de la dernière page).

extractJsonLdBlocks est exporté depuis json-ld-recipe.ts pour être
réutilisé par marmiton.ts sans dupliquer le regex d'extraction des blocs
<script type="application/ld+json">.

Enregistre marmitonAdapter dans registerAllRecipeSources (sources/index.ts)
— contrairement à jsonLdRecipeAdapter lui-même, c'est un adaptateur concret
par site, donc une Source household-toggleable légitime.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* feat(recipes): ajoute un adaptateur RecipeSourceAdapter pour 750g

Suit le même schéma que marmitonAdapter (construit sur jsonLdRecipeAdapter),
avec deux différences propres à 750g.com :

- list() n'a pas d'ItemList JSON-LD à lire sur ses résultats de recherche
  (la recherche du site est un widget client-side) — appelle donc
  directement le endpoint GET que ce widget interroge lui-même en interne
  (un « moteur de réponse IA » qui renvoie un lot de recettes pour une
  requête en texte libre), et scrape les cartes de résultat par regex en
  associant à chaque lien de recette sa dernière image précédente plutôt
  qu'un zip naïf par index (des images décoratives sans carte associée
  existent réellement dans ce fragment). Vérifié en direct : demander une
  « page 2 » revient toujours vide, donc nextCursor vaut toujours null,
  comme theMealDbAdapter.
- parse() ne délègue pas aussi directement à jsonLdRecipeAdapter.parse que
  marmitonAdapter — le générateur JSON-LD de 750g.com a deux bugs réels :
  des caractères de contrôle bruts non échappés dans certaines chaînes JSON
  (~1 recette sur 3 dans un échantillon vérifié en direct, sinon
  JSON.parse échoue et jsonLdRecipeAdapter rapporte à tort « aucun
  Recipe trouvé »), et un texte parfois doublement encodé en entités HTML
  (ex. un vrai « é » devient &amp;eacute; au lieu de &eacute;). Les deux
  sont corrigés en pré/post-traitement autour de la même délégation, pas
  une réimplémentation.

Enregistre sevenFiftyGAdapter dans registerAllRecipeSources
(sources/index.ts), au même titre que marmitonAdapter.

Complète aussi test/sources/sources-index.test.ts, qui ne couvrait encore
que TheMealDB malgré l'ajout de Marmiton dans une PR précédente.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* feat(recipes): ajoute un adaptateur RecipeSourceAdapter pour Manger Bouger

Suit le même schéma que marmitonAdapter/sevenFiftyGAdapter (construit sur
jsonLdRecipeAdapter), avec des différences propres à mangerbouger.fr
(« La Fabrique à Menus », Santé publique France) :

- list() n'utilise pas de JSON-LD du tout — la page de résultats (une app
  Next.js) n'embarque aucun ItemList. Elle est cependant rendue
  côté serveur et expose le même state Redux que le client hydrate, via un
  <script id="__NEXT_DATA__">, qui contient déjà tout ce dont list() a
  besoin (slug/nom/image, pagination). Vérifié en direct : ?query=<texte
  libre> filtre bien côté serveur, et hasMorePages donne un signal de fin
  de pagination plus propre que le 404 de Marmiton ou l'absence de vraie
  pagination de 750g.
- parse() délègue à jsonLdRecipeAdapter mais corrige deux lacunes réelles
  et systématiques de son propre JSON-LD (vérifiées sur 9 recettes,
  72 étapes) : recipeInstructions[].text est un document Slate.js
  sérialisé en JSON (pas du texte) plutôt qu'être aplati ; recipeYield est
  absent partout alors que le nombre de portions existe bien côté site
  (__NEXT_DATA__) — les deux sont corrigés par un patch structuré (parse →
  mutation → réécriture) avant délégation, pas une réimplémentation.

Enregistre mangerBougerAdapter dans registerAllRecipeSources
(sources/index.ts) et complète sources-index.test.ts.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-22 19:01:46 +02:00

355 lines
15 KiB
TypeScript

import type {
ParsedRecipe,
RecipeSourceAdapter,
RecipeSourceListItem,
RecipeSourceListParams,
RecipeSourceListResult,
} from "../lib/recipe-sources/recipe-source-adapter.js";
import {
RecipeSourceFetchError,
RecipeSourceParseError,
} from "../lib/recipe-sources/recipe-source-errors.js";
import { jsonLdRecipeAdapter } from "./json-ld-recipe.js";
const SOURCE_KEY = "mangerBouger";
// "La Fabrique à Menus" — mangerbouger.fr's recipe tool (Santé publique
// France). Its listing page is a Next.js app with no JSON-LD `ItemList` at
// all (unlike marmiton.ts's search page) — but it's server-rendered, and a
// plain GET carries the exact same Redux state the client hydrates from as
// a `__NEXT_DATA__` script tag (see `extractNextData` below), which already
// has everything `list()` needs. Verified live: `?query=<free text>` really
// filters server-side (not just a client-side URL update over an
// already-fetched page), and `page`/`hasMorePages` behave as real,
// consistent pagination — the best-behaved of this adapter family's three
// sources on that front.
const LIST_URL = "https://www.mangerbouger.fr/manger-mieux/la-fabrique-a-menus/recettes";
const DETAIL_BASE_URL = "https://www.mangerbouger.fr/manger-mieux/la-fabrique-a-menus/recettes/";
/** Matches the `<script id="__NEXT_DATA__">…</script>` block every Next.js page ships — the site's own server-rendered hydration data, read instead of scraping HTML for both `list()` (the listing's recipe cards) and `parse()` (backfilling a gap in the detail page's JSON-LD, see {@link extractPortionsFromNextData}). */
const NEXT_DATA_PATTERN = /<script id="__NEXT_DATA__"[^>]*>([\s\S]*?)<\/script>/;
/** The one field of one `list[]` entry `list()` actually reads off the listing page's `__NEXT_DATA__` — that state carries the site's full internal `Recipe` shape (60+ fields: nutriscore, seasons, macros, …), none of which this adapter's contract has anywhere to put. */
interface MangerBougerListEntry {
slug?: string;
name?: string;
image?: string | null;
}
/** The slice of `__NEXT_DATA__` this module reads off the *listing* page. */
interface MangerBougerListPageData {
props?: {
initialState?: {
recipes?: {
list?: MangerBougerListEntry[];
/** Whether a further page exists for the current `page`/`query`/`diet` combination — verified live: an out-of-range page comes back `false` with an empty `list` rather than repeating the last page or erroring, a cleaner end-of-results signal than either `marmiton.ts` (infers it from a 404) or `750g.ts` (this search has no real pagination at all). */
hasMorePages?: boolean;
};
};
};
}
/** The slice of `__NEXT_DATA__` this module reads off a recipe *detail* page — a different shape than the listing page's (`initialState.recipe.recipe`, not `initialState.recipes.list[]`) since it's a different Redux slice entirely. */
interface MangerBougerDetailPageData {
props?: {
initialState?: {
recipe?: {
recipe?: {
portions?: unknown;
};
};
};
};
}
/** Parses the page's `__NEXT_DATA__` block into `T`, or `null` if the block is missing or isn't valid JSON — callers degrade gracefully rather than throw, same as `marmiton.ts`'s "page has no ItemList at all" handling. */
function extractNextData<T>(html: string): T | null {
const match = html.match(NEXT_DATA_PATTERN);
if (!match) return null;
try {
return JSON.parse(match[1] ?? "") as T;
} catch {
return null;
}
}
function detailUrl(slug: string): string {
return `${DETAIL_BASE_URL}${slug}`;
}
/**
* Matches the single `<script type="application/ld+json">…</script>` block
* a mangerbouger.fr recipe *detail* page carries (verified live across a
* sample of 9 recipes — always exactly one, always a bare `Recipe`, never
* an `@graph`) — a much narrower pattern than `json-ld-recipe.ts`'s own
* `JSON_LD_SCRIPT_PATTERN` (no `g` flag: this module only ever needs the
* first/only block, to patch it — see {@link patchRecipeJsonLd}) or
* `750g.ts`'s identically-named private copy (which does its own,
* different, character-level repair over every block on the page).
*/
const JSON_LD_SCRIPT_PATTERN =
/(<script[^>]*type\s*=\s*["']application\/ld\+json["'][^>]*>)([\s\S]*?)(<\/script>)/i;
/** One node of a Slate.js rich-text document — see {@link flattenSlateDocument}. */
interface SlateNode {
type?: string;
text?: string;
children?: SlateNode[];
}
/** Concatenates a run of inline Slate nodes (leaf text, or further-nested inline runs) with no separator — bold/italic/underline marks (the only ones observed) carry no plain-text equivalent and are simply dropped. */
function flattenSlateInline(nodes: SlateNode[]): string {
return nodes
.map((node) =>
typeof node.text === "string"
? node.text
: node.children
? flattenSlateInline(node.children)
: "",
)
.join("");
}
/**
* Flattens a Slate.js document's top-level blocks into one line of plain
* text each — verified live across every recipe step sampled (72 recipes):
* only `paragraph` and `bulleted-list` (of `list-item`s) ever appear as
* block types, so that's all this handles; any other/unrecognized block
* type still degrades reasonably (its own children read as one inline run)
* rather than being dropped outright.
*/
function flattenSlateBlocks(nodes: SlateNode[]): string[] {
const lines: string[] = [];
for (const node of nodes) {
if (node.type === "bulleted-list" && node.children) {
lines.push(...flattenSlateBlocks(node.children));
continue;
}
if (node.type === "list-item" && node.children) {
const text = flattenSlateInline(node.children);
if (text.trim().length > 0) lines.push(`- ${text}`);
continue;
}
if (node.children) {
const text = flattenSlateInline(node.children);
if (text.trim().length > 0) lines.push(text);
continue;
}
if (typeof node.text === "string" && node.text.trim().length > 0) lines.push(node.text);
}
return lines;
}
/**
* Flattens one `HowToStep.text` value into plain text. mangerbouger.fr's
* own JSON-LD embeds this field pre-formatted for its own web app instead
* of as prose: `text` is itself a JSON-serialized Slate.js rich-text
* document (verified live: every one of 72 sampled recipe steps parses as
* one) — handing that straight to `jsonLdRecipeAdapter.parse` would surface
* the raw `[{"type":"paragraph","children":[{"text":"…` blob as a step's
* description, unusable as-is. `json` that doesn't parse as an array (a
* genuinely plain-text step, or some future/different shape) is returned
* unchanged rather than mangled.
*/
function flattenSlateDocument(json: string): string {
let doc: unknown;
try {
doc = JSON.parse(json);
} catch {
return json;
}
if (!Array.isArray(doc)) return json;
return flattenSlateBlocks(doc as SlateNode[]).join("\n");
}
/** The two schema.org `Recipe` fields {@link patchRecipeJsonLd} patches, plus an index signature so every other field survives re-serialization untouched. */
interface JsonLdRecipeLike {
recipeInstructions?: unknown;
recipeYield?: unknown;
[key: string]: unknown;
}
/** One `HowToStep`-shaped entry of `recipeInstructions`, as far as {@link patchRecipeJsonLd} needs to know. */
interface JsonLdHowToStepLike {
text?: unknown;
[key: string]: unknown;
}
/**
* `state.recipe.recipe.portions` from the same detail page's `__NEXT_DATA__`
* — the number `recipeYield` should have been (see {@link patchRecipeJsonLd}),
* read from the site's own internal state rather than left unstated.
*/
function extractPortionsFromNextData(html: string): number | null {
const data = extractNextData<MangerBougerDetailPageData>(html);
const portions = data?.props?.initialState?.recipe?.recipe?.portions;
return typeof portions === "number" ? portions : null;
}
/**
* Repairs the two real gaps verified live in mangerbouger.fr's own
* recipe-detail JSON-LD, then hands the patched HTML to
* `jsonLdRecipeAdapter.parse` unmodified otherwise — same "fix what's
* actually broken, delegate the rest" shape as `750g.ts`'s
* `sanitizeJsonLdBlocks`/`decodeParsedRecipeText`, just structural (parse →
* mutate → re-serialize the one JSON-LD object) rather than textual, since
* both gaps need real understanding of the document, not character-level
* fixups:
*
* - `recipeInstructions[].text` is Slate.js rich text, not prose — flattened
* via {@link flattenSlateDocument}.
* - `recipeYield` is absent on every one of 9 sampled recipes (schema.org
* allows omitting it, and mangerbouger.fr's generator apparently always
* does) even though the site's own internal data has the serving count
* right there — backfilled from `__NEXT_DATA__` via
* {@link extractPortionsFromNextData} rather than left as a needless
* `portions: null` on every single imported recipe.
*
* A missing or malformed JSON-LD block is left completely untouched —
* `jsonLdRecipeAdapter`'s own "no JSON-LD Recipe found"/"malformed block,
* skip it" handling is exactly the right behavior for that, no need to
* duplicate it here.
*/
function patchRecipeJsonLd(html: string): string {
const match = html.match(JSON_LD_SCRIPT_PATTERN);
if (!match) return html;
let recipe: JsonLdRecipeLike;
try {
recipe = JSON.parse(match[2] ?? "{}") as JsonLdRecipeLike;
} catch {
return html;
}
if (Array.isArray(recipe.recipeInstructions)) {
for (const step of recipe.recipeInstructions as JsonLdHowToStepLike[]) {
if (step && typeof step === "object" && typeof step.text === "string") {
step.text = flattenSlateDocument(step.text);
}
}
}
if (recipe.recipeYield === undefined) {
const portions = extractPortionsFromNextData(html);
if (portions !== null) recipe.recipeYield = portions;
}
const patchedJson = JSON.stringify(recipe);
return html.replace(
JSON_LD_SCRIPT_PATTERN,
(_full, openTag: string, _json: string, closeTag: string) =>
`${openTag}${patchedJson}${closeTag}`,
);
}
/**
* Re-labels a `RecipeSourceFetchError`/`RecipeSourceParseError` thrown by
* the generic `jsonLdRecipeAdapter` (`sourceKey` `"jsonLdRecipe"`) as having
* come from this adapter instead (`sourceKey` `"mangerBouger"`) — same
* reasoning as `marmiton.ts`/`750g.ts`'s identically-named helpers.
*/
function rekeySourceError(err: unknown): unknown {
if (err instanceof RecipeSourceFetchError) {
return new RecipeSourceFetchError(SOURCE_KEY, err.message, { cause: err.cause });
}
if (err instanceof RecipeSourceParseError) {
return new RecipeSourceParseError(SOURCE_KEY, err.message, { cause: err.cause });
}
return err;
}
/**
* mangerbouger.fr ("La Fabrique à Menus") — Santé publique France's public
* nutrition site. Unofficial (`official: false`): no published API, same
* reasoning as every other adapter in this family — fetching ordinary pages
* and reading data the site never committed to a stable contract, not a
* maintained endpoint. `fetchDetail` delegates straight to
* `jsonLdRecipeAdapter`; `parse` wraps it with {@link patchRecipeJsonLd}
* (see that function's doc comment for the two real gaps it fixes).
* `list()` doesn't use JSON-LD at all — see `LIST_URL`'s doc comment.
*/
export const mangerBougerAdapter: RecipeSourceAdapter<{ html: string; url: string }> = {
key: SOURCE_KEY,
name: "Manger Bouger",
official: false,
// Chemin fixe (pas d'icône versionnée/hashée comme sur d'autres sources
// de cette famille) — répond correctement sans paramètre supplémentaire.
iconUrl: "https://www.mangerbouger.fr/manger-mieux/la-fabrique-a-menus/favicon.ico",
// Le contenu de mangerbouger.fr (noms, ingrédients, instructions) est en
// français — détermine contre quel modèle/locale d'étiquettes
// d'ingrédients translateRecipe (recipe-translation.ts) résout les
// recettes de cette source lors d'une prévisualisation/d'un import.
locale: "fr",
async list(params: RecipeSourceListParams): Promise<RecipeSourceListResult> {
try {
const page = params.cursor ? Number(params.cursor) : 1;
const query = params.query ?? "";
const listUrl = `${LIST_URL}?diet=ALL&page=${page}&query=${encodeURIComponent(query)}`;
let response: Response;
try {
response = await fetch(listUrl);
} catch (cause) {
throw new RecipeSourceFetchError(SOURCE_KEY, `Network error listing recipes (${listUrl})`, {
cause,
});
}
if (!response.ok) {
throw new RecipeSourceFetchError(
SOURCE_KEY,
`mangerbouger.fr responded ${response.status} (${listUrl})`,
);
}
const html = await response.text();
const data = extractNextData<MangerBougerListPageData>(html);
const state = data?.props?.initialState?.recipes;
const items: RecipeSourceListItem[] = (state?.list ?? [])
.filter((entry): entry is MangerBougerListEntry & { slug: string; name: string } =>
Boolean(entry.slug && entry.name),
)
.map((entry) => ({
externalId: detailUrl(entry.slug),
title: entry.name,
picture: entry.image ?? null,
url: detailUrl(entry.slug),
}));
return { items, nextCursor: state?.hasMorePages ? String(page + 1) : null };
} catch (err) {
// Rethrown as-is (already keyed "mangerBouger" by whichever branch
// above threw it) — this adapter's only caller (`sources.service.ts`)
// already handles/logs failures centrally; this method just isn't
// allowed a bare `async` body without a try/catch per the repo's
// convention. Same reasoning as `marmiton.ts`/`750g.ts`.
throw err;
}
},
// `externalId` est directement l'URL canonique de la recette sur
// mangerbouger.fr (renvoyée telle quelle par `list()` ci-dessus) — même
// convention que `jsonLdRecipeAdapter.fetchDetail`, à qui cette méthode
// délègue entièrement : la réparation du JSON-LD (voir
// `patchRecipeJsonLd`) n'a lieu qu'à l'étape `parse()`, pas ici.
async fetchDetail(externalId: string): Promise<{ html: string; url: string }> {
try {
return await jsonLdRecipeAdapter.fetchDetail(externalId);
} catch (err) {
// Pas un simple re-throw : `rekeySourceError` est le traitement utile
// que ce point d'appel doit faire de l'erreur (relabelliser sa
// `sourceKey`), conformément à la convention await/try-catch du repo.
throw rekeySourceError(err);
}
},
parse(raw: { html: string; url: string }): ParsedRecipe {
try {
const patchedHtml = patchRecipeJsonLd(raw.html);
return jsonLdRecipeAdapter.parse({ html: patchedHtml, url: raw.url });
} catch (err) {
throw rekeySourceError(err);
}
},
};