Animatrixx
VisitSolo engineer — backend, infrastructure, frontendAnime & Manga Platform · Self-hosted Infrastructure
Animatrixx is an anime streaming and manga reading platform. The manga side is fully operational: a catalogue ingested from MangaDex with cover and banner art reconciled from AniList, debounced full-text search backed by Elasticsearch, tag filtering, threaded comments, a per-user recency rail and trending and popular rankings, and a fullscreen reader with vertical-scroll and page modes, keyboard shortcuts and auto-hiding chrome. The anime side ships the complete viewing interface — player controls, quality and playback-speed selection, seek and volume with mute interlock, episode navigation and YouTube-style keybindings — with the streaming backend as the current build phase rather than a finished layer.
The reason this project earns a case study is the half users never see. It runs on nine containers I orchestrate and administer: Django 5 under uWSGI behind nginx, MariaDB, Redis, RabbitMQ, Celery worker and beat, Elasticsearch, and a Prometheus, Grafana, Loki and Promtail observability stack scraping application metrics every four seconds and shipping structured logs off the box.
I chose to self-host all of it. A managed platform would have been faster and I would have learned considerably less — most of what is written below is a consequence of having to run the thing rather than deploy it.
- Open the live product
- Private repository — walkthrough on request
nginx
static from volume, uwsgi_pass to app
Django + uWSGI
DRF, token auth, nested routers
Catalogue path
Elasticsearch search, Redis-cached rails
Reader path
image proxy, 7-day immutable cache
Async + observability plane
Celery, RabbitMQ, Prometheus, Grafana, Loki, Promtail
01 / 09
Problem
Manga sources do not want to be aggregated. MangaDex serves page images from ephemeral, hotlink-protected nodes: there are no stable URLs, you must call an allocator to be assigned a node and compose URLs yourself from a returned hash and filename list, and those nodes reject any request whose Referer is not mangadex.org. A correctly-composed URL still returns 403 from a browser on another domain.
The metadata is split across two sources with no shared identifier. MangaDex has complete chapter data and poor cover art; AniList has excellent cover and banner art and no chapter pages. There is no ISBN, no cross-reference — the only available join is the title, and MangaDex titles arrive as a language-keyed dictionary full of romanisation noise, punctuation and season markers that AniList’s search will not match.
Both APIs rate-limit on separate, undocumented budgets, and a catalogue crawl makes thousands of sequential calls. A naive loop dies partway and leaves the database half-populated; a naive retry storm gets the address banned. And the ranking features users expect — continue reading, trending, popular — are ranked lists that change on every page view, which is a GROUP BY over a view-events table on every homepage load if you build them the obvious way.
02 / 09
Goal
Infrastructure goal
Run the whole system myself, on servers I administer, with enough observability to diagnose a problem from metrics and logs rather than by guessing. Concretely: application metrics scraped on a short interval, structured logs shipped off the host, and a dashboard I can open before I open a shell.
Ingestion goal
A crawl that survives contact with two hostile rate limits:
- Never abort the whole run for a single failed item — degrade to skipping that item.
- Differentiate the retry strategy by failure class, rather than retrying everything identically.
- Stay under the sustained-rate radar deliberately, with jitter, instead of going as fast as possible and reacting to bans.
- Make re-runs idempotent, so an interrupted crawl can simply be run again.
Read-path goal
Keep ranked, personalised lists off the relational database. No aggregate query on the homepage request path, and no analytics write blocking a user’s page load.
03 / 09
Architecture
Nine containers on one bridge network, plus a Next.js frontend deployed separately. The shape below is the production composition; the observability plane is drawn detached because it observes the system rather than serving it.
Edge — nginx
nginx (unprivileged, alpine) terminates the request, serves collected static assets straight from a shared volume so they never touch an application worker, and forwards everything else to Django over the uwsgi protocol rather than HTTP.
Application — Django 5 and DRF under uWSGI
Django 5.0.3 with Django REST Framework, running under uWSGI in master mode with threads enabled. Token authentication, page-number pagination, eleven ordered middleware, and nested routers from drf-nested-routers so chapters and comments are properly parented under a manga rather than filtered by query parameter. Three stacked auth backends, including a custom one that accepts either a username or an email in the same field.
State — MariaDB, Redis, Elasticsearch, S3
MariaDB holds the relational truth: seven models across four apps, with the manga-to-tag many-to-many, self-referential threaded comments, and a per-user reading-history row constrained unique on user and manga. Redis serves double duty — a django-redis cache for the non-personalised rails, and a separate raw client holding sorted sets and lists for rankings. Elasticsearch mirrors the searchable fields as a read replica so full-text queries never reach MariaDB. S3 in ap-south-1 holds site assets.
Asynchronous work — Celery on RabbitMQ
A Celery worker and a beat scheduler, brokered over AMQP with results persisted back to MariaDB. View tracking is fired from the detail endpoint with delay() so a user’s page load never waits on analytics writes; welcome emails render an HTML template with a plain-text alternative and go out through SMTP off the request path.
Read path — the image proxy
The one piece of infrastructure that exists purely because an upstream refuses to cooperate. A Django view validates the requested host against an allowlist, re-requests the image with a spoofed Referer, and streams it back with a seven-day immutable cache header — so the browser only ever talks to my origin and the ephemeral upstream node stays an implementation detail.
Observability — Prometheus, Grafana, Loki, Promtail
django-prometheus exposes application and database-query metrics; Prometheus scrapes them every four seconds. Promtail tails the Django log and ships it to Loki, which stores it on a TSDB schema. Grafana sits on top. This is the layer I would have been given at a larger company and had to build here, and building it is the reason I can read one.
Configurations for distributed tracing and long-term metrics storage exist in the repository but are not wired into the running composition. They are staged work, not shipped work, and I would rather say so than let a directory listing imply otherwise.
Screens pending — live product linked above
04 / 09
Key features
Manga catalogue and reader
Ingested catalogue, tag filtering, chapter list with sort, fullscreen reader with vertical and page modes, keyboard shortcuts, auto-hiding chrome and a page-grid navigator
The engineering underneathChapter numbers arrive as free text, so ordering is a database-level cast to integer on a derived column rather than the model’s lexicographic default
Full-text search
Debounced search across titles, descriptions and tags with its own pagination bounds
The engineering underneathElasticsearch as a read replica of the searchable fields, so query load never lands on MariaDB. 500ms debounce on the client so keystrokes do not each become a query
Personalised and ranked rails
Continue reading, trending and popular, on the homepage
The engineering underneathMaintained incrementally in Redis rather than recomputed. Recency uses a capped-LRU list idiom that self-bounds at ten entries with no cleanup job
Image proxy
Serves upstream page images through my own origin
The engineering underneathHost-allowlisted, Referer-spoofing relay with a seven-day immutable cache header — the only way those images can reach a browser on this domain at all
Threaded comments
Nested replies with edit and delete restricted to the author
The engineering underneathSelf-referential foreign key with a recursive serializer, and the one query in the project hardened against N+1 with select_related and prefetch_related
Auth, including Google
Username-or-email and password, or Google OAuth
The engineering underneathCustom authentication backend resolving either identifier; OAuth client credentials read from the database at request time rather than baked into the application
Anime viewing interface
Full player chrome — seek, volume with mute interlock, quality and speed selection, fullscreen, episode navigation, auto-hide, keybindings
The engineering underneathInterface complete and interactive; the streaming backend is the current build phase, so the player runs against a placeholder source rather than a manifest
05 / 09
Technical decisions
Four decisions I would defend, and one I would reverse. The reversal is the more useful conversation.
Four retry strategies, chosen by failure class.
Treating every failure identically is how a crawler either gives up too early or gets banned. So each class gets its own response: rate limiting and dropped connections retry on truncated exponential backoff plus random jitter — the jitter matters because a synchronous crawler retrying on exact powers of two self-synchronises into bursts; timeouts retry on plain exponential backoff; non-retryable HTTP errors bail immediately rather than burning the retry budget; and the second upstream, with its own separate limit, gets its own fixed-interval loop. Between items there is a deliberate jittered politeness delay, because staying under the sustained-rate threshold on purpose is cheaper than reacting to bans.
Skip the item, never the run.
If a chapter’s images cannot be resolved after five attempts, the crawl logs it and continues to the next chapter rather than persisting a broken row or aborting. Upserts are keyed on the upstream identifier pair, so a re-run repairs what was skipped instead of duplicating what succeeded. Combined, those two properties mean the recovery procedure for a failed crawl is to run it again — which is the only recovery procedure I will reliably follow.
Reprojecting Redis ordering back into SQL.
The recency rail needs rows in insertion order, and a WHERE id IN (...) query explicitly does not guarantee it. The obvious fix is to fetch and re-sort in Python. Instead the Redis list order is compiled into a CASE WHEN ladder passed to order_by, so the database returns rows already in the right order and the application does no sorting at all. It is the sharpest few lines in the backend and the one I would put on a whiteboard.
Self-hosting the observability stack.
Prometheus scraping every four seconds, Promtail tailing into Loki, Grafana on top, with the database engine itself wrapped so query-level metrics are exposed alongside application metrics. A hosted APM would have been an afternoon. Building it taught me what a scrape interval costs, why log shipping is a separate concern from log writing, and how to read a metrics plane — which is the transferable skill, and the reason this was worth the extra week.
Handing chapter images to the reader through localStorage. I would reverse this.
The chapter list writes the page-URL array into localStorage and the reader reads it back, which avoids a second request and works perfectly from a click. It also means the reader cannot be refreshed, deep-linked or shared — arriving directly renders an error — and because the list endpoint serialises every field, a long series ships every chapter’s full image array to fetch one. The correct design is a per-chapter endpoint and a trimmed list serializer. This is the clearest example on this page of a decision that was locally convenient and architecturally wrong, and I would rather name it than have someone find it.
06 / 09
Challenges
Joining two catalogues with no shared key.
The only join available was the title, across two sources that spell things differently. The normaliser is deliberately lossy — strip everything non-alphanumeric, keep the first four words — because maximising recall matters more than precision when a miss simply means no cover art. Once resolved, the discovered second identifier is stored, so the fuzzy match runs once per title and never again.
Chapter numbers that are not numbers.
Upstream chapter numbers include values like 10.5 and Extra, along with empty and null, so the field is text and the model’s default ordering sorts chapter 10 before chapter 2. Solved twice, in two places: a database-level cast to integer on an annotated column for the API ordering, and again on the client because the chapter list also supports a user-toggled sort direction that cannot re-query.
A producer and consumer that disagreed on a key.
The view-tracking task writes a per-day trending key; the homepage reads a weekly one; nothing rolls the daily keys up. The result is that the trending rail returns an empty list and the daily sets accumulate unread. The architecture is right and the aggregation job is missing — the fix is one scheduled union of the last seven days. I found this by reading the code path rather than from an alert, which is itself the finding: I have metrics on request rates and none on whether a feature returns data.
Secrets baked into the image.
The Dockerfile copies the environment file into the image and the docker ignore file is empty, so credentials are present in every layer. This is the most serious defect in the project, it is entirely mine, and the remediation is not just a config change — the keys have to be rotated because they are also in the build history. Naming it here is deliberate: I would rather be the person who found it than the person who shipped it and did not notice.
No tests, on a system with real users.
There is no automated test suite in either repository — the test files are the Django scaffolding with nothing in them. MoneyDock has 272 tests because Animatrixx taught me what their absence costs: every one of the defects above was found by reading code or by a user hitting it, which is the slowest and most expensive way to find anything.
Turning off the type checker to ship.
The frontend build ignores TypeScript errors, and that flag is load-bearing: the manga detail page renders six fields the serializer never returns, so they are undefined in production. Disabling the check did not make the problem go away, it made it invisible — and the honest version of this is that I chose shipping over correctness and then stopped being able to see the difference.
07 / 09
Lessons learned
Build the observability plane once, and you can read anyone’s.
Standing up Prometheus, Loki, Promtail and Grafana myself is why a metrics dashboard is now a tool rather than a wall of charts. It is the single most transferable thing this project gave me, and I would not have got it from a hosted agent and a default dashboard.
Instrument features, not just infrastructure.
I had four-second scrape intervals on request rates and no way to notice that a homepage rail had been returning an empty list. Infrastructure metrics tell you the service is up. They say nothing about whether it is doing its job, and I now treat those as two separate questions.
Suppressing a warning is a decision with a cost.
Ignoring build-time type errors let me ship and hid a live data-contract break. The lesson was not “types are good” — it was that turning off a check converts a known problem into an unknown one, and unknown problems are the expensive kind.
Idempotence is the recovery plan you will actually use.
Upserting on a stable upstream key means the answer to a half-finished crawl is to run it again. Every recovery procedure more complex than that is a procedure I would have got wrong at the moment I needed it.
Convenient at the call site, wrong at the boundary.
The localStorage handoff saved a request and cost addressability, shareability and refresh. I now ask what a design forbids, not only what it enables — a URL that cannot be shared is a product decision disguised as an implementation detail.
This project is the reason the next one had tests.
Almost every discipline in MoneyDock — the test-gated deploy, the pure functions, the timeouts on every outbound call, the hardening pass with each fix pinned by a test — is a direct response to something that went wrong here. Taken as a pair, that progression is the most useful thing either project says about me.
08 / 09
Future improvements
Ordered by severity rather than by appeal. The first three are correctness and security work, not features.
Rotate every credential and rebuild the image with a populated docker ignore file, so secrets stop living in build layers and history. Highest priority, and not optional.
Ship the streaming backend behind the completed player interface — source resolution, an HLS manifest and a segment-serving path — so the anime half runs on real media rather than a placeholder.
Add the scheduled rollup that unions the daily trending keys into the weekly one, and put a TTL on the dailies so they stop accumulating unread.
Make the reader addressable: a per-chapter images endpoint, a trimmed list serializer that stops shipping every chapter’s pages, and real previous and next chapter links instead of the current dead buttons.
Add lazy loading and windowing to the reader. Every page of a chapter currently mounts at once against a proxy that buffers whole images in a worker, so a long chapter fires dozens of concurrent requests on open.
Turn the type checker back on and fix what it finds, starting with the six fields the manga detail page reads and the serializer never sends.
Fix the persistence path for reading progress — two endpoints currently fail on a type mismatch, which is why the reader never reports progress at all.
Add throttling. The rate-limit decorator on login is installed and commented out, so there is no brute-force protection on the authentication endpoint.
Standard Django production hardening: TLS redirect, HSTS, secure cookie flags, a proxy SSL header, and log rotation — the log currently runs at debug level to an unrotated file, which captures every query forever.
A test suite. Starting with the ingestion pipeline and the auth endpoints, because those are where a regression is both most likely and least visible.
09 / 09
Technologies used
In the codebase
- Python 3.10
- Django 5
- Django REST Framework
- Celery
- RabbitMQ
- Redis
- MariaDB
- Elasticsearch
- Docker Compose
- nginx
- uWSGI
- Prometheus
- Grafana
- Loki
- Promtail
- AWS S3
- Next.js 15
- React 19
- TypeScript
- Tailwind CSS
Engineering practices the build depends on
- Container orchestration
- Linux server administration
- Reverse-proxy configuration
- Distributed task queues
- Cache design
- Web scraping at scale
- Backoff and jitter
- Observability engineering
- OAuth integration