Chains are real Xray hops: a chain-<code> inbound + outbound + routing rule on the
entry node and a relay-<code> client on the exit node, reconciled by sync_node with
config test, port check, rollback on a failed restart and a per-node lock.
Admin API gets chains CRUD/probe/check, node latency, audit log (middleware, no request
bodies), users list with subscription counters and a csv export that neutralises formulas.
Each user gets a short ref_code (backfilled lazily for pre-existing
accounts too) and a shareable t.me/<bot>?start=ref_<code> link, new
"Пригласить друга" menu item shows it plus how many referrals actually
converted and any bonus days waiting to be applied.
Reward fires once, on the referred user's first subscription of any
kind (free, gift, or paid) — not on signup, so an unconverted click
never pays out. Both sides get REFERRAL_BONUS_DAYS (config.py/.env,
default 3): the referrer's day count comes from settings.py's live-read
pattern, same as prices/HWID, so it's tunable without a restart even
before a panel UI exists for it. Bonus extends an active subscription
directly if the recipient has one, otherwise accumulates in
bonus_days_pending and gets folded into whichever subscription they
create next (redeemed automatically inside create_subscription, one
choke point regardless of which of bot.py's several call sites created
it — free trial, gift code, paid, or admin grant).
Guards: no self-referral, referrer must exist, first-touch attribution
only (a second ?start=ref_ link never overwrites it), and only takes
for genuinely new accounts (no existing subscriptions) — attaching a
referrer to an already-active user was never the intent.
Tested two ways, matching this repo's usual db.py-can-be-imported-
standalone / bot.py-needs-a-workaround split: 16 checks against a real
isolated sqlite db for the db.py logic (attribution, both reward paths,
double-reward guard, pending-bonus fold-in), then 9 more through an
actual `import bot` — aiogram/fastapi now have Python 3.14 wheels so
this imported for real rather than needing AST-extraction, modulo one
old blocker (xray_manager still imports the Unix-only fcntl for its
file lock) worked around with a tiny fake fcntl module in sys.modules,
same spirit as the fcntl shim already used elsewhere in this project's
history. Real start_deeplink and cb_referral calls, get_me() mocked to
avoid a live Telegram API call.
Not done: admin-panel UI toggle for REFERRAL_ENABLED/REFERRAL_BONUS_DAYS
(currently .env-only, like several other business tunables were before
they got a settings-page treatment) and a docs-tab writeup — happy to
add both if wanted, scoped this pass to the mechanic itself.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
New lens this pass: read every admin-mutating route for input
validation, not just auth (already audited that separately). Found
POST /admin/api/nodes taking `code` straight from the request body with
zero checks — reachable for real from admin.html's "Вручную" add-node
tab (nm-code is a free-text field), not just a theoretical API-only
path. code is the table's PRIMARY KEY and gets embedded directly into
every /admin/api/nodes/{code}/... URL afterward.
Concretely: an empty code, one containing a slash, or one that
collides with an existing node would all previously either succeed
into a node the UI can no longer address by its own generated URLs, or
crash with a raw unhandled sqlite3.IntegrityError / ValueError instead
of a real error message. None of this needed a live server to
reproduce — it's pure input handling.
Traced kind="external" nodes first before touching anything near
them, since add_client_to_node/remove_client_from_node only branch on
"local"/"managed" with no external case — worth being sure that's the
intentional "this node's clients are managed outside the panel, we
just reference a fixed shared_uuid" design (confirmed via links.py's
own use of shared_uuid) and not an actual bug before writing validation
around it.
NODE_CODE_RE (same style as the existing HWID_RE): letters/digits/-/_,
1-32 chars — covers "de1", the "n"+hex(4) auto-generated codes, and any
reasonable manual name, rejects anything that would break URL routing
or silently create an unreachable node. Non-empty label. Port coerced
and range-checked (1-65535) instead of a bare int() that throws on
garbage input. Duplicate code now raises a ValueError from
db.create_node (pre-checked via get_node(), same pattern create_admin
already uses for duplicate usernames — not a bolted-on try/except
IntegrityError) which the route turns into a real 400.
Deliberately scoped to creation only — code isn't in admin_update_node's
editable set, so there's no separate update-path gap to also close.
Verification: db.create_node's duplicate guard tested directly against
a real sqlite db (fresh code succeeds, immediate duplicate attempt
raises and leaves the original untouched, a second distinct code still
works). NODE_CODE_RE run through 12 cases — valid codes including the
real auto-generated shape, and the specific invalid ones that matter
(slash, space, unicode, empty, over-length, exactly-at-the-length-
limit). AST-extracted the whole updated admin_create_node() route (still
can't import api.py) and drove it through a fake db/webhooks/
HTTPException with 9 cases covering every rejection branch, the happy
path (including that node.added still fires with the right payload),
and the duplicate-code path specifically, confirming the ValueError
from db.py correctly surfaces as an HTTP 400 rather than an unhandled
exception.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
From the original night's low-priority backlog item ("user on hold
status") — the only lever admin had for cutting a customer's access was
Revoke, which is permanent: the subscription's remaining days are just
gone, and restoring access means manually granting a brand-new one and
eyeballing how many days to give back. No way to say "block this for a
few days, then give the exact remaining time back."
db.py: new held_at column on subscriptions (same ALTER-TABLE migration
pattern as every other column added this week). hold_subscription()
sets it, guarded to only fire on a subscription that's currently active,
not already held, not expired — returns False instead of silently
no-opping so the caller can tell holding didn't happen. resume_subscription()
shifts expires_at forward by exactly how long it was held (now - held_at)
and clears held_at, so a subscription paused for 3 days comes back with
3 days added, not 3 days lost.
The part that actually mattered for correctness: list_active_subscriptions()
now also requires held_at IS NULL. This function is what xray_manager's
periodic sync (every 90s) uses to decide which clients belong in Xray's
config — without this exclusion, holding a subscription would look like
it worked for about 90 seconds and then the next sync would silently
re-add the client, since the row still has active=1 and a future
expires_at. Found this by actually tracing sync_from_db()/sync_all()
before writing the hold logic, not after debugging a live failure.
api.py: POST .../hold and .../resume routes, mirroring the existing
revoke route (fetch the sub, touch the node's xray client immediately
rather than waiting for the next periodic sync, same as revoke already
does). _days_left() now takes the whole subscription row instead of just
expires_at, so it can use held_at as the reference point instead of "now"
for a held subscription — otherwise the admin UI would show the days
counter silently ticking down while the customer isn't even able to use
the service.
admin.html: Пауза/Возобновить buttons next to Отозвать in both the main
Подписки table and the per-user card, a "на паузе" badge, and a doc-block
explaining the hold-vs-revoke distinction. Also fixed a latent race while
touching this code: the old inline revoke handler in the user card fired
openUserCard() immediately alongside revokeSub() without waiting for it,
so the card could refresh before the revoke's own API call had finished;
switched to .then() so hold/resume/revoke all correctly wait for the
action before refreshing the card.
Verification: db.py has no fastapi/aiogram dependency so this was fully
testable locally, unlike most of tonight's api.py/bot.py-touching work.
16 checks against a real isolated sqlite db: hold/resume round-trip,
the exclude-from-active-list behavior the xray sync depends on, the
exact hours-shift math (simulated a 5h hold by rewriting held_at
directly, verified the resumed expires_at landed within 6 minutes of
the expected shift), and edge cases — double-hold, double-resume,
holding an expired or already-revoked subscription, nonexistent uuid.
AST-extracted the updated _days_left() out of api.py (still can't
import the module directly) and ran it against hand-built held/active
subscription dicts. Added the same hold/resume sequence to the existing
CI "TOTP/backup/reorder" step and ran that step's exact full script
locally end to end before committing — all six of its sections pass
together, not just the new one in isolation.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Neither endpoint had any brute-force protection — TOTP codes are only
6 digits (1M combinations) and HMAC-SHA1 verification is cheap, so an
unthrottled /admin/api/login/totp is a realistic brute-force target
within a pending token's 5-minute window. Password login had the same
gap.
DB-backed (new login_attempts table), not in-memory — this matters
now that mbs-api runs multiple worker processes (see 11c75c1): an
in-process counter would let an attacker split requests across
workers and bypass it entirely, same class of mistake as an
unsynchronized in-memory cache. Keyed by client IP (nginx already
sets X-Real-IP on every proxied request, install.sh has always done
this).
10 failed attempts / 15min for password, 10 / 5min for TOTP codes,
counted per-IP per-kind. Successful login clears that IP's recent
failures. Old rows pruned in the existing 90s periodic_sync cleanup
alongside sessions and pending_totp.
Verified: threshold counting, per-IP isolation, per-kind isolation
(password vs totp tracked separately), clear-on-success, and the
age-based cleanup only removing rows older than the cutoff.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Matches Remnawave's security-first positioning (passkeys/OAuth there)
with the more universal standard instead — RFC 6238 TOTP, works with
Google Authenticator/Authy/1Password/anything. Implemented from
scratch on stdlib only (hashlib/hmac/struct/base64) — no new
dependency — and verified against the official RFC 4226 HOTP test
vectors (all 10 pass exactly) before wiring it into any auth path.
Per-admin, optional: setup shows the secret + otpauth:// URI (no QR
render, just copyable text — didn't want to fake a QR library),
confirmed by entering a real code before it's persisted. Login is now
two-step when 2FA is on: password first (returns a short-lived
pending_token instead of a session if totp_secret is set), then a
second call with the pending_token + code creates the real session.
pending_totp rows are single-use and expire in 5 minutes, cleaned up
alongside the existing admin_sessions cleanup in periodic_sync.
Disabling 2FA requires re-entering the current password.
Verified end-to-end: enable/disable round trip, pending-token
resolve+single-use+expiry, code verification against the stored
secret, wrong-code and wrong-password rejection — on top of the raw
HOTP correctness check.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Matches Marzban's multi-admin (WIP there) and closes a real gap vs
both. New 'admins' table (username + PBKDF2-SHA256 password hash,
200k iterations, random salt per account, stdlib hashlib/hmac only —
no new dependency), admin_sessions now tracks which admin is logged
in. Existing installs aren't broken: on first run, if no admins exist
yet, a default 'admin' account is seeded from the current
ADMIN_PANEL_PASSWORD — old password keeps working under username
'admin', pre-filled on the login screen.
Admin management lives in Settings: list, add (username + password,
min 8 chars), remove. Can't delete the last remaining admin or your
own currently-logged-in account. Sidebar now shows who's logged in.
Verified end-to-end: bootstrap, correct/wrong/nonexistent login,
session->admin resolution, last-admin-delete protection, duplicate
username rejection, add/remove round trip, and that identical
passwords hash to different values (unique salt) but both verify.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Every backup restore was leaving a .before-restore-<timestamp> safety
copy of both mbs.db and .env with no cleanup — would accumulate
forever on a panel that restores regularly. Now keeps the 5 most
recent and prunes the rest right after each restore. Verified: 8
fake copies pruned down to exactly the 5 newest, oldest-first.
admin_sessions rows also never got deleted once expired (only ever
filtered out of queries via expires_at>?) — same shape of problem.
Piggybacks on the existing 90s periodic_sync heartbeat in bot.py
instead of adding a new one.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Matches Remnawave/Marzban's 'Host Sorting via Web UI' — was the one
concrete UI gap flagged in the comparison research. Nodes get a
persistent sort_order column (backfilled once from the existing
de1-first/created_at order on migration, verified idempotent — re-running
_migrate() on an already-migrated db does not reshuffle it back).
Plain HTML5 drag-and-drop (dragstart/dragover/drop), no library —
matches the project's no-frameworks admin.html. Drop reorders the
in-memory list optimistically, re-renders immediately, then persists
via POST /admin/api/nodes/reorder; a failed save reloads from the
server instead of leaving the UI out of sync with the db.
db.reorder_nodes() rejects any list that doesn't contain exactly the
current set of node codes (no silent drops or duplicates). New nodes
append at the end (MAX(sort_order)+1) instead of jumping to the front.
Verified with a full test: initial order, reorder, two rejected
malformed reorders, a new node appending at the end, and migration
re-run not touching an already-backfilled order.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
payments: _grant_paid_subscription now validates plan/node exist before
marking a payment paid instead of after (was leaving charged-but-ungranted
payments with no error trail); mark_payment_paid is now a single atomic
UPDATE ... WHERE status='pending' instead of check-then-act, closing a
double-grant race between webhooks and the periodic reconciler; yookassa
webhook now re-verifies payment status server-side via the API instead of
trusting the posted body (platega already had HMAC verification).
hwid: 'user["hwid_limit"] or FALLBACK' treated an explicit 0 (admin fully
blocking a user) as unset — now an explicit None check. Device count-check
and insert are now one atomic transaction (db.add_device_if_under_limit)
instead of two raceable statements.
perf: payment webhooks and _grant_paid_subscription's SSH/HTTP calls now
run via asyncio.to_thread instead of blocking the event loop; same for
bot.py's periodic_sync/reconcile_pending_payments and the manual admin
sync button. Admin endpoints (traffic/subscriptions/payments/gift-codes/
user-card) now resolve node labels from one db.list_nodes() call instead
of a fresh db.get_node() per row. revoke/reset-traffic use a direct PK
lookup instead of scanning up to 5000 rows. Dashboard now asks the API
for 8 rows instead of fetching 200 and slicing client-side.
security: mbs.db (and -wal/-shm) now chmod 600 right after creation —
it held session tokens and subscription bearer tokens world-readable
by default. Node SSH connections now pin host keys via a persisted
known_hosts file (TOFU) instead of accepting any key on every connection.
delete_node now refuses to delete a node with active subscriptions
instead of silently orphaning their xray clients.
deadcode: removed unused xray_manager.list_client_ids and admin.html's
superseded staggerReveal (rows animate via rowAttr() inline now).
Also guards gift-code redemption against a plan/node deleted after the
code was created (was an unhandled KeyError/TypeError crash).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>