Files
multica/server/internal/integrations/lark/http_client_test.go
beast 60048172a7 fix(lark): ingest inbound images and videos as chat attachments (MUL-4934) (#5580)
* fix: ingest feishu media as chat attachments

* fix: ingest feishu post embedded media

* fix(lark): make inbound media retries safe

* fix lark media resource limit

* fix(lark): move inbound media off ack path

* fix(channel): make inbound media runs durable

* fix(channel): close enqueue-vs-append race on media deferral

EnqueueChatTask read the session-wide media deadline in one statement and
sealed the input batch in a later one. Under READ COMMITTED a media message
committing between the two got sealed into a task the deadline read had
already decided was 'queued', so the daemon could claim it before its
attachment bound — the agent received the bare placeholder, and the later
media-ready promotion was a no-op against a non-deferred task.

After the seal, re-derive the deferral from the sealed batch itself in the
same transaction (DeferChatTaskForSealedPendingMedia): if any sealed message
still carries an unexpired media marker, flip the task to deferred with
fire_at aligned to the latest marker. The existing post-commit promote fence
already covers the opposite direction (marker cleared mid-transaction).

Adds a deterministic regression test that injects the media append between
the deadline read and the seal via a wrapped pgx.Tx.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channel): keep committed chat task out of enqueue error path

The post-commit media-ready fence returned its error from EnqueueChatTask
even though the deferred task was already durably committed. The router
flush treats any enqueue error as "no task exists": it clears the typing
indicator and logs an enqueue failure while the run still happens at its
fire_at deadline. Log the fence failure instead — the claim-path deferred
promoter re-queues the task regardless.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channel): cap global media resolution concurrency

Media jobs were serialized per session but unbounded across sessions: a
burst could open arbitrarily many concurrent 45s Lark downloads, and each
unknown-length upload may buffer up to the 100 MiB resource cap in memory.
Gate resolveAndBindMedia behind a global slot semaphore (default 8,
RouterConfig.MediaConcurrency). Per-session ordering is unchanged; on
shutdown a job cancelled while waiting for a slot proceeds straight to the
bounded DB finalize so marker clearing stays prompt. Also document that the
per-message media budget spans queue/slot waits (it must match the
persisted fire_at) and why timed-out uploads cannot leak unbounded orphans.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(chat): keep channel-sealed user messages on task cancel

Sealing the channel input batch stamps task_id onto channel user
messages, which exposed them to the cancel draft-restore path: an
empty-transcript cancel would DeleteUserChatMessageByTask the sealed
Feishu/Slack messages and detach their attachments. Those messages are
the durable record of what the platform sender wrote — the sender has
no Multica composer to restore a draft into.

Gate the restore-delete on ChatSessionHasChannelBinding in both the
synchronous finalize and the deferred finalize (the latter covers
markers left by an older replica during a rolling deploy); a bound
session now settles as "Stopped." instead.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channel): skip the media pipeline for messages without media

Every inbound message on a Media-enabled platform persisted a 45s
media deadline and queued a resolution job, so a plain text message
could wait behind the global media semaphore (its task deferred while
other sessions download 100 MiB videos) and a crash between append and
clear delayed a pure-text run to the full 45s fallback.

Add MediaResolver.HasMedia — a pure in-memory probe the Router calls
on the ACK path — and only persist the deadline / enqueue the job when
the message actually references platform media. The Feishu resolver
decodes the already-received payload and reports standalone image or
video keys and post-embedded img/media spans.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(chat): gate cancel restore on immutable channel provenance

The previous guard keyed the cancel restore-delete off
ChatSessionHasChannelBinding, but a binding only proves routing exists
right now: archiving a session and rebinding an installation both
delete the binding while preserving chat history, so a still-cancellable
sealed task could again restore-delete the original inbound messages.

Persist provenance on the message instead: migration 203 adds
chat_message.channel_ingested, stamped inside the channel append
transaction and never mutated, and both cancel finalize paths now gate
on TaskHasChannelIngestedMessages over the task's sealed batch. The
binding-existence query is removed. Regression tests cover ingest ->
archive/unbind -> cancel for a queued and a started task.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channel): reclaim media uploads that never gain an attachment row

Deadline expiry dropped already-resolved refs and a BindMedia failure
was log-only, leaving uploaded objects with no attachment row and no
reclaim path — the dedup mark commits with the message before media
runs, so a redelivery is dropped as a duplicate and never re-resolves
(and thus never overwrites) those keys, and workspace/session deletion
only enumerates the attachment table.

Add MediaResolver.DiscardMedia — a best-effort delete by StorageKey —
and call it from both failure paths in resolveAndBindMedia. The Feishu
resolver forwards to the storage backend's Delete. Tests cover a
partial upload discarded at the deadline, discard on bind failure, and
key-level deletion in the resolver.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(server): refresh comments stale after detached media ingestion

Channel tasks now seal a self-owned input batch, media ingestion is no
longer out of scope for the flattener, and MediaRefs are filled by the
detached resolver after append rather than by feishuChannel pre-engine.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(chat): stop keying channel empty-completion silence off chat_input_task_id

Sealing gave channel tasks a self input-owner, which broke
writeChatCompletionOutcome's discriminator: it treated any owned task
as direct, so an empty channel completion wrote the no_response
fallback row and the outbound patcher — which forwards any non-empty
chat:done content verbatim — pushed the English fallback body to
Feishu/Slack, violating the MUL-4351 contract.

Silence is now decided by the immutable channel_ingested provenance of
the task's input batch, looked up by the batch OWNER id
(chat_input_task_id): auto-retry clones inherit the owner while their
sealed messages stay tagged with the parent's id, so keying off the
task's own id would misread a channel retry as direct. The cancel-path
provenance gates switch to the same owner key via chatInputOwnerID.
chat_input_task_id is back to meaning only "input batch owner".

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(migrations): renumber to 203/204 after upstream took 202

Upstream main merged 202_runtime_profile_add_qwen while this branch
held 202/203, tripping TestMigrationNumericPrefixesStayUniqueAfterLegacySet
on the CI merge tree. channel_media_pending becomes 203 and
channel_ingested becomes 204; no content changes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channels): gate outbound delivery on channel provenance, not owner

Merging main brought #5645 (keep direct chat replies in Multica),
whose outbound gate assumed channel tasks leave chat_input_task_id
NULL. Sealed channel tasks own an input batch too, so on the merge
tree every channel reply and failure notice was classified as direct
and silently dropped — agents stopped replying in Feishu/Slack.

Both outbound gates now call engine.TaskInputIsChannelIngested: a NULL
owner keeps #5645's deliver-by-default for pre-sealing tasks, an owned
batch delivers only when it carries the immutable channel_ingested
stamp (keyed by the owner id, so auto-retry clones inherit the
verdict). Direct replies stay in Multica; sealed channel replies reach
the platform. Tests cover both directions on both platforms.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channel): discard media orphans on a fresh context, after finalize

DiscardMedia shared finalizeCtx with BindMedia, so a bind that failed
because the finalize deadline expired handed the storage deletes an
already-dead context — the compensation silently no-opped and the
orphans leaked anyway. The deadline path also ran S3 deletes before
the marker clear, eating the same 5s budget the user-facing
bind/promotion needed.

Collect the refs from both failure paths, run bind + promotion on the
finalize budget first, then delete on a fresh discard context. The
bind-failure test now pins that discard receives a live context.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channel): compensate result-uncertain media uploads and commits

The compensation protocol treated "the call returned an error" as "the
side effect did not happen", which is wrong in both directions across
the result-uncertain windows:

- An upload error can follow a server-side write (lost response,
  deadline mid-write). The attempted key never reached the router, so
  nothing could reclaim it and dedup guarantees no re-resolve. The
  resolver now idempotently deletes the deterministic key on a fresh
  budget right at the failure site.

- A commit error is not a rollback guarantee: a lost ack can report
  failure after Postgres durably committed the attachment rows, and
  the router's discard would then delete objects those rows reference.
  BindMediaRefs now converges the ambiguity on a fresh budget — any of
  the batch's URLs present proves the atomic commit landed (bind
  reports success); none proves the rollback (discard stays safe); a
  failed verification returns ErrMediaBindResultUnknown and the router
  keeps the uploads, preferring a rare orphan over a broken attachment.

Fault-injection coverage: an upload error deletes the attempted key; a
lost-ack commit keeps the bound attachment and reports success; a
verified rollback stays a discardable error; the router keeps uploads
on the unknown-outcome sentinel.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(migrations): renumber to 207/208 after upstream took 203-206

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* docs(channel): note DiscardMedia self-invocation and the unknown-outcome skip

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(migrations): renumber to 212/213 after upstream took 207-211

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat(channel): replace inline media compensation with an intent ledger and reconciler

Inline best-effort compensation cannot answer "did my side effect
happen?" at the moment it needs the answer — the DELETE/PUT reordering
and the empty-read-vs-in-flight-COMMIT gaps were both instances of the
same two-system atomicity problem. Persist the intent instead and let
an asynchronous reconciler settle it:

- channel_media_pending_object (migration 214; claim index 215 as its
  own single-statement CONCURRENTLY migration): a state machine row
  ('pending' -> 'deleting') with lease, attempt, and backoff columns.
- The resolver upserts the row BEFORE each PUT, state-guarded so a key
  the reconciler owns is never resurrected (the resource is skipped).
  ObjectURL is a pure function of configuration, so the row carries the
  attachment URL pre-upload.
- BindMediaRefs deletes the batch's rows INSIDE the attachment-insert
  transaction: commit landed <=> intents gone, atomically, so an
  ambiguous COMMIT never needs adjudication. A key already claimed to
  'deleting' is skipped (placeholder stays).
- Nothing is ever deleted inline. The reconciler — an independent
  worker so storage latency cannot starve other sweepers — claims due
  rows ('pending' past the settle delay, or expired leases) under a
  fresh lease, checks for a durable attachment reference only AFTER the
  claim (race-free: bind can no longer succeed on the key), deletes
  unreferenced objects outside any transaction, and backs off failed
  deletes with attempt-based retry. Crash windows converge for free.
- The settle delay is a fixed constant carrying NO correctness weight;
  invariant tests pin it at >=10x every pipeline budget. Metrics cover
  deletes, referenced clears, delete failures, and ledger backlog.

Removed: MediaResolver.DiscardMedia, ErrMediaBindResultUnknown, the
post-commit verification, and both router discard branches.

Tests: intent-before-upload ordering; upload error leaves the row and
deletes nothing; bind-wins vs reconciler-wins on the same key; lost-ack
and rolled-back commit injections (intent cleared iff the attachment
landed); reconciler three-state settle; expired-lease reclaim; delete
failure backoff and retry; settle invariants.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channel): never build or sweep the media reconciler without storage

store is nil when S3 is unconfigured AND the local upload dir fails to
initialize, but the reconciler was constructed unconditionally and
main only gates the goroutine on the reconciler pointer — the first
unreferenced ledger row (rows can pre-exist from a boot where storage
worked) would nil-pointer panic a bare goroutine and take down the
process.

Construct the reconciler only when a storage backend exists, and guard
RunOnce defensively: with no deleter it skips the sweep without
claiming, so rows are not stranded in 'deleting' until lease expiry.
Test covers the pre-existing-row + missing-storage boot.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(migrations): renumber to 213-216 after upstream took 212

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(channel): remove the dead pre-resolved MediaRefs ingress path

lark.InboundMessage.MediaRefs and the resolver's early-returns for
pre-populated refs were vestiges of the pre-detached synchronous design
— no producer fills them before the router anymore. Worse, the intent
ledger made the path actively misleading: refs arriving without ledger
rows would be silently skipped at bind (with a log blaming the
reconciler), contradicting the field's "already persisted" contract.

Delete the field, its channelMessageFromLark mapping, and both
early-returns; channel.InboundMessage.MediaRefs is now documented as
what it actually is — ResolveMedia's output channel, always empty on
ingress, attachable only through a claimed ledger intent.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channel): enforce workspace tenancy on every ledger query

The intent-ledger upsert's conflict branch guarded only on state, so a
cross-workspace storage_key collision could rewrite the row's
workspace/message/url ownership; release and delete keyed on
(storage_key, lease_token) alone. The derived key embeds the workspace
UUID so none of this is reachable today — but tenancy must be enforced
by the workspace column in every query, never derived from the key
string (MUL-3515 rule, restated in this PR's review).

The upsert now updates only within the same workspace (a cross-tenant
conflict updates nothing, returns no row, and the resolver skips the
upload — the fail-safe direction), and release/delete take
(workspace_id, storage_key, lease_token). Tests pin that a foreign
workspace can neither steal, release, nor delete a row.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(migrations): build the ledger primary key via a concurrent index

storage_key TEXT PRIMARY KEY created its unique index implicitly at
CREATE TABLE, against the repo convention that every migration index —
including a new table's unique index — is built CONCURRENTLY in its
own single-statement migration (the exact three-step pattern
client_usage_daily shipped in 207-209). The table now declares
storage_key NOT NULL, 216 builds the unique index concurrently, and
217 attaches the primary key USING INDEX; the claim index moves to
218. ON CONFLICT (storage_key) still resolves against the constraint,
and the full down/up round-trip is verified.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channel): bound each reconciler object delete with its own timeout

DeleteObject ran on the worker-lifetime context and the SDK's default
HTTP client has no overall request timeout, so one black-holed
connection would wedge the sequential sweep loop — and with it every
later batch and the backlog gauge — forever; a single-replica
deployment has no other worker to reclaim the lease. Each delete now
gets a 30s timeout (well under the 2min lease), and a timed-out delete
takes the existing release/backoff path. Covered by a blocking-deleter
test with an injectable timeout.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channel): anchor the media deadline to the DB clock and bound queue waits by it

Two deadline gaps from review:

- The persisted marker was an application-clock timestamp compared
  against SQL now() everywhere it is read, so a skewed app node could
  shrink the fallback window and hand the agent a placeholder before
  the resolver's local budget ended. The append transaction now anchors
  a relative budget (MediaPendingSeconds) with now() + make_interval,
  writer and readers sharing one clock; the local resolve budget stays
  monotonic app-side. A DB test pins that the remaining budget measured
  by the DB clock equals the requested one.

- enqueueMedia's waits (per-session order, global slot) only watched
  shutdown, so in a burst an already-expired job kept its goroutine and
  payload until it reached the front. Both waits now also watch the
  message's deadline; on expiry the job skips the resolver entirely and
  runs only the empty finalize (marker clear + promotion), which also
  unblocks the session's later messages. Covered by a queued-expiry
  test that finalizes while the only slot is deterministically held.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(migrations): renumber to 216-221 after upstream took 213-215

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channel): start the local media budget before the append transaction

The DB anchors the durable fallback at insert-time now(), but the
local monotonic budget started only after AppendMessage returned — so
the resolver outlived the fallback by the append/commit latency, a
window where the deferred task is already claimable while the resolver
still runs and the agent reads a placeholder that binds moments later.
Capture the local deadline before calling AppendMessage, restoring the
ordering local-gives-up <= durable-fallback-fires. A slow-append test
pins that the resolver's context deadline is measured from the
pre-append instant.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* chore(migrations): renumber to 224-229 after upstream took 216-223

Verified against the merged tree: the numeric-prefix uniqueness test
passes and the full migration set applies cleanly from scratch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channel): heartbeat the reconciler lease per row

One claim covers up to 50 rows under a single 2-minute lease, but the
batch is processed sequentially and each delete may run its full 30s
timeout — a few stalled deletes could outlive the lease mid-batch,
letting another replica reclaim the tail: duplicate concurrent
deletes, inflated attempt/backoff on rows whose owner was alive, and
skewed metrics.

The lease is now renewed before EACH row's settle work, so it only
ever needs to cover one row's worst case (invariant-tested: lease >=
2x the per-delete timeout). A renewal that matches no row means
another worker reclaimed it after a genuine expiry — the row is
skipped, leaving the new owner's state untouched. Test simulates a
mid-batch reclaim and pins the skip.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channel): dedup post media resources and make local writes atomic

A rich post may reference the same image_key/file_key in several spans.
The object key derives from (message, type, key), so duplicates uploaded
to the SAME key twice: LocalStorage.UploadStream truncated the
destination up front and removed it outright on a copy error, so a
second failing attempt destroyed the object the first success had
produced — leaving an attachment row pointing at nothing. A second
succeeding attempt instead produced two attachment rows for one object.

Collapse duplicate spans by (fetch type, platform key) before the
upload loop, and write local uploads through a temp file renamed into
place so a failed write can only discard its own temp file. Tests cover
a duplicated span uploading once and a failed re-upload leaving the
previous object intact.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channel): fence late-materializing PUTs with a tombstone schedule

A DELETE cannot be ordered against a PUT the client already abandoned:
the store may materialize the object after the delete completes. The
reconciler cleared the ledger row right after deleting, so such an
object had no row and nothing to reclaim it — which made the settle
delay the de-facto correctness barrier for the PUT/DELETE race, exactly
what the design says it must not be.

The row is now kept as a tombstone ('tombstoned' state, migration 226's
CHECK) and re-deleted on a widening schedule (15m, 1h, 6h, 24h, the
pass index carried in last_error), so a late materialization is
reclaimed by a later pass; only after the schedule is exhausted is the
row dropped. Claim, heartbeat, lease, and tenancy predicates are
unchanged — a tombstone is claimed exactly like any other due row. A
separate gauge reports tombstones so they cannot be mistaken for a
backlog of objects awaiting reclaim, and the header comment now states
precisely what state fences (bind/commit) versus what the schedule
fences (late PUTs).

Tests: the reviewer's interleaving — DELETE completes, the abandoned PUT
materializes right after, and the object is gone by the end of the
schedule — plus a full schedule walk asserting the object is counted
once and the row clears at the end.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channel): tombstones must re-delete, not re-ask the reference question

A tombstone revisit ran the same reference check as a first settle, so
an attachment carrying the same URL — a re-ingested copy of the object —
sent the row down the "referenced, keep it" branch: the object was kept
and the row cleared, abandoning the re-delete schedule that fences the
ORIGINAL object against an abandoned PUT. A tombstone has already been
judged unreferenced and deleted; it exists only to re-delete whatever
materializes later, so it now goes straight to the delete + schedule
tail (extracted as settleDeletedObject, shared with the first settle).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(channel): keep the tombstone schedule position in its own column

The re-delete pass index was encoded into last_error, which the failure
path also writes: one failed re-delete erased the position and restarted
the walk. A store failing intermittently could therefore keep a tombstone
alive indefinitely — every recovery would resume at pass 1 and the row
would never reach the end of the schedule to be dropped.

tombstone_pass is now its own column (the table is introduced in this PR,
so migration 226 carries it), advanced only by a successful delete, and
the tombstone write clears the now-stale last_error. Test walks the
schedule across a failed re-delete and asserts it resumes rather than
restarts, and that the row still terminates.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(lark): derive media object keys per chat message

The object key was derived from the platform message alone, so a second
ingest of the same Feishu message reused the first ingest's ledger row.
That row can be a tombstone (up to ~31h while the re-delete schedule
runs), and the intent upsert refuses anything that has left 'pending', so
the second ingest skipped the upload and silently produced a placeholder
with no attachment. A re-ingest is reachable: the inbound dedup claim is
reclaimable once 60s stale and the dedup row is only vacuumed after 24h.

Keying on the chat message the object will attach to keeps the two
ingests independent, and nothing leaks: each one's objects are covered by
its own ledger row.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(storage): route both local upload paths through one atomic write

UploadStream wrote through a temp file and renamed into place, but the
buffered Upload path still truncated the destination up front — the
destructive shape the stream path exists to avoid, one caller away from
coming back. Both now share writeAtomic, which also restores the 0644 the
direct write used (CreateTemp makes files 0600).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(channel): gofmt the media-pending append fields

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(storage): keep the local upload chmod best-effort

The rename-into-place rewrite made a failed chmod fail the whole upload.
CreateTemp's 0600 has to be widened to the 0644 the direct write used, but
an upload dir on a mount that ignores chmod (SMB/NFS/FUSE) accepted the
old direct write fine — turning those deployments' uploads into hard
errors would be a regression for a cosmetic property. Log and continue.

Tests pin 0644 on both upload paths, and that a failed buffered upload
leaves no temp litter and no damage to a previous object.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(channel): never re-delete an object an attachment references

The tombstone pass skipped the reference check and deleted unconditionally,
so a durable attachment carrying that URL lost the only object it can read
— the dangling attachment the intent ledger exists to prevent, and the
opposite of the posture every other path here takes ("a reclaimable orphan
beats a broken attachment").

The check now runs on every pass. A positive result on a tombstone is
unreachable by design — keys are per (chat message, resource) and a bind
cannot attach a key that has left 'pending' — so reaching it means an
invariant broke: keep the object, clear the row, log it, and count it on
a dedicated reconciler_tombstone_referenced_total counter. The test's
contract is flipped to assert the referenced object survives.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(storage): make the local staging file reclaimable after a crash

os.CreateTemp's random suffix meant a crash between the staging write and
the rename left a file nothing could name: the ledger records only the
final storage key, and DeleteObject removed only the object and its
sidecar. Each leftover can approach the 100 MiB resource cap and they
accumulate without bound.

The staging path is now derived from the object key, so DeleteObject
removes it alongside the object — which makes the media reconciler reclaim
it too, since the intent row is written before the upload. Opening it 0644
directly also drops the chmod the previous commit had to make best-effort.
Both read paths refuse the staging name (keys come from the request URL,
and a half-written body should not be readable); a user-supplied ".tmp"
extension is unaffected, since object keys are generated and never
dot-prefixed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(channel): renumber the media migrations after merging main

main took 224 (agent_task_session_rollout_missing), so the ledger group
moves to 225-230 and the cross-references inside the table migration follow.
main's CompleteTask also grew a sessionRolloutMissing parameter; the three
call sites this PR added to chat_input_ownership_test.go pass false.

Verified the way the numbering is meant to be verified: full migration set
applied from scratch on the merged tree, and the whole server suite run
against that database.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(channel): let Postgres compute every reconciler deadline

The reconciler built settle cutoffs, lease expiry, backoff and re-delete
times from the process clock and compared them against the database's
now(). A replica whose clock had drifted would therefore settle rows whose
upload was still in flight (the object is deleted and the bind then refuses
to attach — media silently lost), hand out leases that are born expired
(rows churn between workers, attempt/backoff inflate), or compress the
tombstone schedule that fences a late-materializing PUT.

The four settle queries now take durations and derive their timestamps from
now(), so every replica reads one clock. The parameter types are the guard:
an app-side timestamp can no longer be passed. Test asserts the persisted
lease, backoff and re-delete deadlines all track the database's now().

The generated code also picks up main's new agent_task_queue column in the
two RETURNING task.* queries this PR adds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* chore(lark): drop unrelated gofmt-only churn from this PR

Six files carried whitespace/comment-reformatting with no functional
change, unrelated to the inbound media pipeline. Reverting them to the
base revision keeps the diff focused on the feature (75 -> 69 files):

  server/internal/service/empty_claim_cache.go
  server/internal/integrations/lark/markdown_detect.go
  server/internal/integrations/lark/ws_chunk_assembler.go
  server/internal/integrations/lark/ws_chunk_assembler_test.go
  server/internal/integrations/lark/ws_frame_test.go
  server/internal/integrations/lark/registration_test.go

Verified: `git diff -w` against these files was already empty, so no
behavior is affected. go vet clean; tests covering these files pass.

Co-authored-by: multica-agent <github@multica.ai>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Bohan-J <bohan@devv.ai>
Co-authored-by: multica-agent <github@multica.ai>
2026-07-27 15:44:39 +08:00

1507 lines
53 KiB
Go
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
package lark
import (
"context"
"encoding/json"
"errors"
"io"
"net/http"
"net/http/httptest"
"strconv"
"strings"
"sync/atomic"
"testing"
"time"
)
// larkFakeServer is a tiny in-memory stand-in for the Lark Open
// Platform. Tests register handlers per path; the server panics if a
// path is hit without a registration (a missed assertion is louder
// than a 404).
//
// The handler shape mirrors http.HandlerFunc so each test can encode
// its own response without inheriting boilerplate.
type larkFakeServer struct {
t *testing.T
mux *http.ServeMux
srv *httptest.Server
tokenN atomic.Int32
sendN atomic.Int32
patchN atomic.Int32
bindN atomic.Int32
reactN atomic.Int32
delRN atomic.Int32
authObs atomic.Value // last Authorization header seen across all paths
}
func newLarkFake(t *testing.T) *larkFakeServer {
t.Helper()
f := &larkFakeServer{t: t, mux: http.NewServeMux()}
f.srv = httptest.NewServer(f)
t.Cleanup(f.srv.Close)
return f
}
func (f *larkFakeServer) URL() string { return f.srv.URL }
func (f *larkFakeServer) ServeHTTP(w http.ResponseWriter, r *http.Request) {
if a := r.Header.Get("Authorization"); a != "" {
f.authObs.Store(a)
}
f.mux.ServeHTTP(w, r)
}
func (f *larkFakeServer) lastAuth() string {
v, _ := f.authObs.Load().(string)
return v
}
// stubToken installs a token endpoint that returns the supplied token
// with the supplied expire (seconds) and counts hits.
func (f *larkFakeServer) stubToken(token string, expireSec int64) {
f.mux.HandleFunc("/open-apis/auth/v3/tenant_access_token/internal", func(w http.ResponseWriter, r *http.Request) {
f.tokenN.Add(1)
if r.Method != http.MethodPost {
f.t.Errorf("token: want POST, got %s", r.Method)
}
var body map[string]string
if err := json.NewDecoder(r.Body).Decode(&body); err != nil {
f.t.Errorf("token: decode body: %v", err)
}
if body["app_id"] == "" || body["app_secret"] == "" {
f.t.Errorf("token: missing app credentials: %v", body)
}
writeJSON(w, map[string]any{
"code": 0,
"msg": "ok",
"tenant_access_token": token,
"expire": expireSec,
})
})
}
// stubTokenError installs a token endpoint returning a Lark-style
// error code (non-zero `code` with HTTP 200).
func (f *larkFakeServer) stubTokenError(code int, msg string) {
f.mux.HandleFunc("/open-apis/auth/v3/tenant_access_token/internal", func(w http.ResponseWriter, r *http.Request) {
f.tokenN.Add(1)
writeJSON(w, map[string]any{"code": code, "msg": msg})
})
}
// stubSend installs the IM-send endpoint. resp is the response body
// (typically the standard {code, msg, data:{message_id}} shape).
func (f *larkFakeServer) stubSend(resp map[string]any, verify func(r *http.Request, body map[string]string)) {
f.mux.HandleFunc("/open-apis/im/v1/messages", func(w http.ResponseWriter, r *http.Request) {
f.sendN.Add(1)
if r.Method != http.MethodPost {
f.t.Errorf("send: want POST, got %s", r.Method)
}
var body map[string]string
if err := json.NewDecoder(r.Body).Decode(&body); err != nil {
f.t.Errorf("send: decode body: %v", err)
}
if verify != nil {
verify(r, body)
}
writeJSON(w, resp)
})
}
// stubReply installs the IM-reply endpoint
// (POST /open-apis/im/v1/messages/<id>/reply), used by the thread-reply
// path. Body is decoded as map[string]any because reply_in_thread is a
// bool. Register stubToken + stubReply (and not stubSend / stubPatch) in
// a reply test, since they share the /messages/ prefix.
func (f *larkFakeServer) stubReply(resp map[string]any, verify func(r *http.Request, id string, body map[string]any)) {
const prefix = "/open-apis/im/v1/messages/"
f.mux.HandleFunc(prefix, func(w http.ResponseWriter, r *http.Request) {
if r.Method != http.MethodPost || !strings.HasSuffix(r.URL.Path, "/reply") {
f.t.Errorf("reply: want POST .../reply, got %s %s", r.Method, r.URL.Path)
return
}
f.sendN.Add(1)
id := strings.TrimSuffix(strings.TrimPrefix(r.URL.Path, prefix), "/reply")
if id == "" {
f.t.Errorf("reply: missing message id")
}
var body map[string]any
if err := json.NewDecoder(r.Body).Decode(&body); err != nil {
f.t.Errorf("reply: decode body: %v", err)
}
if verify != nil {
verify(r, id, body)
}
writeJSON(w, resp)
})
}
// stubPatch installs the IM-patch endpoint. The Lark route is
// /open-apis/im/v1/messages/<id>; ServeMux uses prefix matching when
// we register the parent path explicitly. We register the parent
// SEND path above already, so the patch path needs the full prefix.
func (f *larkFakeServer) stubPatch(resp map[string]any, verify func(r *http.Request, id string, body map[string]string)) {
const prefix = "/open-apis/im/v1/messages/"
f.mux.HandleFunc(prefix, func(w http.ResponseWriter, r *http.Request) {
if r.Method != http.MethodPatch {
f.t.Errorf("patch: want PATCH, got %s", r.Method)
}
id := strings.TrimPrefix(r.URL.Path, prefix)
if id == "" {
f.t.Errorf("patch: missing message id")
}
f.patchN.Add(1)
var body map[string]string
if err := json.NewDecoder(r.Body).Decode(&body); err != nil {
f.t.Errorf("patch: decode body: %v", err)
}
if verify != nil {
verify(r, id, body)
}
writeJSON(w, resp)
})
}
// stubReaction installs the IM-reaction-create endpoint.
func (f *larkFakeServer) stubReaction(resp map[string]any, verify func(r *http.Request, id string, body map[string]any)) {
const suffix = "/reactions"
f.mux.HandleFunc("/open-apis/im/v1/messages/", func(w http.ResponseWriter, r *http.Request) {
if !strings.HasSuffix(r.URL.Path, suffix) {
return // let other handlers match
}
if r.Method != http.MethodPost {
f.t.Errorf("reaction: want POST, got %s", r.Method)
}
f.reactN.Add(1)
rawID := strings.TrimSuffix(strings.TrimPrefix(r.URL.Path, "/open-apis/im/v1/messages/"), suffix)
if rawID == "" {
f.t.Errorf("reaction: missing message id")
}
var body map[string]any
if err := json.NewDecoder(r.Body).Decode(&body); err != nil {
f.t.Errorf("reaction: decode body: %v", err)
}
if verify != nil {
verify(r, rawID, body)
}
writeJSON(w, resp)
})
}
// stubReactionDelete installs the IM-reaction-delete endpoint.
func (f *larkFakeServer) stubReactionDelete(resp map[string]any, verify func(r *http.Request, msgID string, reactionID string)) {
const prefix = "/open-apis/im/v1/messages/"
f.mux.HandleFunc(prefix, func(w http.ResponseWriter, r *http.Request) {
if r.Method != http.MethodDelete {
return // let other handlers match
}
rest := strings.TrimPrefix(r.URL.Path, prefix)
parts := strings.Split(rest, "/reactions/")
if len(parts) != 2 {
return // not a delete path
}
f.delRN.Add(1)
if parts[0] == "" {
f.t.Errorf("reaction delete: missing message id")
}
if parts[1] == "" {
f.t.Errorf("reaction delete: missing reaction id")
}
if verify != nil {
verify(r, parts[0], parts[1])
}
writeJSON(w, resp)
})
}
func writeJSON(w http.ResponseWriter, body any) {
w.Header().Set("Content-Type", "application/json")
_ = json.NewEncoder(w).Encode(body)
}
// newTestClient returns an httpAPIClient pointed at the fake server,
// using the supplied clock so token expiry can be controlled
// deterministically.
func newTestClient(fake *larkFakeServer, now func() time.Time) *httpAPIClient {
c := NewHTTPAPIClient(HTTPClientConfig{
BaseURL: fake.URL(),
Now: now,
}).(*httpAPIClient)
return c
}
func testCreds() InstallationCredentials {
return InstallationCredentials{AppID: "cli_app_xx", AppSecret: "secret_xx"}
}
func TestHTTPClient_IsConfigured(t *testing.T) {
c := NewHTTPAPIClient(HTTPClientConfig{})
if !c.IsConfigured() {
t.Fatalf("real client must report IsConfigured()=true")
}
}
func TestHTTPClient_DownloadMessageResource(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_resource", 7200)
fake.mux.HandleFunc("/open-apis/im/v1/messages/om_1/resources/img_1", func(w http.ResponseWriter, r *http.Request) {
if r.Method != http.MethodGet {
t.Errorf("resource: method = %s, want GET", r.Method)
}
if got := r.URL.Query().Get("type"); got != "image" {
t.Errorf("resource type = %q, want image", got)
}
if got := r.Header.Get("Authorization"); got != "Bearer tok_resource" {
t.Errorf("auth = %q", got)
}
w.Header().Set("Content-Type", "image/png")
w.Header().Set("Content-Disposition", `attachment; filename="shot.png"`)
_, _ = w.Write([]byte{1, 2, 3})
})
c := newTestClient(fake, time.Now)
got, err := c.DownloadMessageResource(context.Background(), testCreds(), DownloadResourceParams{
MessageID: "om_1",
FileKey: "img_1",
Type: "image",
})
if err != nil {
t.Fatalf("DownloadMessageResource: %v", err)
}
if string(got.Data) != string([]byte{1, 2, 3}) || got.ContentType != "image/png" ||
got.Filename != "shot.png" || got.SizeBytes != 3 {
t.Fatalf("downloaded resource wrong: %+v", got)
}
}
func TestHTTPClient_DownloadMessageResourceUsesResourceTimeout(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_resource", 7200)
fake.mux.HandleFunc("/open-apis/im/v1/messages/om_video/resources/file_video", func(w http.ResponseWriter, r *http.Request) {
if got := r.URL.Query().Get("type"); got != "file" {
t.Errorf("resource type = %q, want file", got)
}
time.Sleep(80 * time.Millisecond)
w.Header().Set("Content-Type", "video/mp4")
w.Header().Set("Content-Disposition", `attachment; filename="clip.mp4"`)
_, _ = w.Write([]byte("slow-video"))
})
c := NewHTTPAPIClient(HTTPClientConfig{
BaseURL: fake.URL(),
HTTPClient: &http.Client{Timeout: 20 * time.Millisecond},
ResourceDownloadTimeout: 500 * time.Millisecond,
Now: time.Now,
}).(*httpAPIClient)
got, err := c.DownloadMessageResource(context.Background(), testCreds(), DownloadResourceParams{
MessageID: "om_video",
FileKey: "file_video",
Type: "file",
})
if err != nil {
t.Fatalf("DownloadMessageResource slow video: %v", err)
}
if string(got.Data) != "slow-video" || got.ContentType != "video/mp4" || got.Filename != "clip.mp4" {
t.Fatalf("downloaded slow video wrong: %+v", got)
}
}
func TestHTTPClient_DownloadMessageResourceExceedingTimeoutIsCancelled(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_resource", 7200)
fake.mux.HandleFunc("/open-apis/im/v1/messages/om_timeout/resources/file_timeout", func(w http.ResponseWriter, r *http.Request) {
select {
case <-r.Context().Done():
return
case <-time.After(500 * time.Millisecond):
_, _ = w.Write([]byte("too late"))
}
})
c := NewHTTPAPIClient(HTTPClientConfig{
BaseURL: fake.URL(),
ResourceDownloadTimeout: 20 * time.Millisecond,
Now: time.Now,
}).(*httpAPIClient)
_, err := c.DownloadMessageResource(context.Background(), testCreds(), DownloadResourceParams{
MessageID: "om_timeout",
FileKey: "file_timeout",
Type: "file",
})
if !errors.Is(err, context.DeadlineExceeded) {
t.Fatalf("timeout error = %v, want context deadline exceeded", err)
}
}
func TestHTTPClient_DownloadMessageResourceRejectsDeclaredOversize(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_resource", 7200)
fake.mux.HandleFunc("/open-apis/im/v1/messages/om_large/resources/file_large", func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "video/mp4")
w.Header().Set("Content-Length", strconv.FormatInt(maxMessageResourceBytes+1, 10))
w.WriteHeader(http.StatusOK)
})
c := newTestClient(fake, time.Now)
_, err := c.DownloadMessageResourceStream(context.Background(), testCreds(), DownloadResourceParams{
MessageID: "om_large",
FileKey: "file_large",
Type: "file",
})
if err == nil || !strings.Contains(err.Error(), "resource exceeds") {
t.Fatalf("declared oversize error = %v", err)
}
}
func TestHTTPClient_DownloadMessageResourceAllowsDeclaredFeishuLimit(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_resource", 7200)
fake.mux.HandleFunc("/open-apis/im/v1/messages/om_limit/resources/file_limit", func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "video/mp4")
w.Header().Set("Content-Length", strconv.FormatInt(maxMessageResourceBytes, 10))
w.WriteHeader(http.StatusOK)
})
c := newTestClient(fake, time.Now)
got, err := c.DownloadMessageResourceStream(context.Background(), testCreds(), DownloadResourceParams{
MessageID: "om_limit",
FileKey: "file_limit",
Type: "file",
})
if err != nil {
t.Fatalf("declared Feishu-limit resource should be accepted: %v", err)
}
got.Body.Close()
if got.SizeBytes != maxMessageResourceBytes {
t.Fatalf("SizeBytes = %d, want %d", got.SizeBytes, maxMessageResourceBytes)
}
}
func TestHTTPClient_DownloadMessageResourceAllowsAbovePreviousLocalLimit(t *testing.T) {
const previousLocalLimit = 20 << 20
fake := newLarkFake(t)
fake.stubToken("tok_resource", 7200)
fake.mux.HandleFunc("/open-apis/im/v1/messages/om_above_previous/resources/file_above_previous", func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "video/mp4")
w.Header().Set("Content-Length", strconv.FormatInt(previousLocalLimit+1, 10))
w.WriteHeader(http.StatusOK)
})
c := newTestClient(fake, time.Now)
got, err := c.DownloadMessageResourceStream(context.Background(), testCreds(), DownloadResourceParams{
MessageID: "om_above_previous",
FileKey: "file_above_previous",
Type: "file",
})
if err != nil {
t.Fatalf("resource above previous 20MiB local limit should be accepted: %v", err)
}
got.Body.Close()
if got.SizeBytes != previousLocalLimit+1 {
t.Fatalf("SizeBytes = %d, want %d", got.SizeBytes, previousLocalLimit+1)
}
}
func TestMaxBytesReadCloserEnforcesUnknownLengthBoundary(t *testing.T) {
t.Run("exact limit", func(t *testing.T) {
body := &maxBytesReadCloser{r: io.NopCloser(strings.NewReader("abc")), remaining: 3}
got, err := io.ReadAll(body)
if err != nil || string(got) != "abc" {
t.Fatalf("exact limit: body=%q err=%v", got, err)
}
})
t.Run("one byte over", func(t *testing.T) {
body := &maxBytesReadCloser{r: io.NopCloser(strings.NewReader("abcd")), remaining: 3}
got, err := io.ReadAll(body)
if err == nil || !strings.Contains(err.Error(), "resource exceeds") || string(got) != "abc" {
t.Fatalf("overflow: body=%q err=%v", got, err)
}
})
}
func TestHTTPClient_DownloadMessageResourceBusinessError(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_resource", 7200)
fake.mux.HandleFunc("/open-apis/im/v1/messages/om_1/resources/img_1", func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "application/json")
writeJSON(w, map[string]any{"code": 234003, "msg": "File not in msg"})
})
c := newTestClient(fake, time.Now)
_, err := c.DownloadMessageResource(context.Background(), testCreds(), DownloadResourceParams{
MessageID: "om_1",
FileKey: "img_1",
Type: "image",
})
if err == nil || !strings.Contains(err.Error(), "234003") {
t.Fatalf("expected APIError with code, got %v", err)
}
}
// TestHTTPClient_StubReportsNotConfigured pins that the stub never
// claims wired outbound — handlers gate install / management UI on
// this signal.
func TestHTTPClient_StubReportsNotConfigured(t *testing.T) {
s := NewStubAPIClient(nil)
if s.IsConfigured() {
t.Errorf("stub IsConfigured must be false")
}
}
// TestHTTPClient_SendInteractiveCard_DefaultRendererBodyHasUpdateMulti
// is the send-side half of the must-fix wire check: when the Patcher
// uses NewDefaultRenderer to produce a card and ships it via
// SendInteractiveCard, the actual HTTP body Lark receives must carry
// config.update_multi=true so the card is patchable downstream.
// Without this, the first send succeeds but every subsequent patch
// silently no-ops on Lark's side while local DB status still flips.
func TestHTTPClient_SendInteractiveCard_DefaultRendererBodyHasUpdateMulti(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_um_send", 7200)
var capturedContent string
fake.stubSend(
map[string]any{"code": 0, "data": map[string]string{"message_id": "om_send_um"}},
func(_ *http.Request, body map[string]string) {
capturedContent = body["content"]
},
)
r := NewDefaultRenderer()
render, err := r.Render(RenderInput{Kind: CardKindThinking, AgentName: "TestAgent"})
if err != nil {
t.Fatalf("render: %v", err)
}
c := newTestClient(fake, time.Now)
if _, err := c.SendInteractiveCard(context.Background(), SendCardParams{
InstallationID: testCreds(),
ChatID: ChatID("oc_send_um"),
CardJSON: render.JSON,
}); err != nil {
t.Fatalf("send: %v", err)
}
assertCardContentHasUpdateMulti(t, capturedContent)
}
// TestHTTPClient_PatchInteractiveCard_DefaultRendererBodyHasUpdateMulti
// is the patch-side half of the same wire check. Every PatchCardParams
// the Patcher produces goes through the default renderer; the body
// shipped over PATCH /open-apis/im/v1/messages/:id must still carry
// update_multi=true, otherwise Lark refuses to apply the patch to a
// card that was sent with update_multi=true (the two ends must agree).
func TestHTTPClient_PatchInteractiveCard_DefaultRendererBodyHasUpdateMulti(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_um_patch", 7200)
var capturedContent string
fake.stubPatch(
map[string]any{"code": 0, "msg": "ok"},
func(_ *http.Request, _ string, body map[string]string) {
capturedContent = body["content"]
},
)
r := NewDefaultRenderer()
render, err := r.Render(RenderInput{Kind: CardKindRunning, AgentName: "TestAgent"})
if err != nil {
t.Fatalf("render: %v", err)
}
c := newTestClient(fake, time.Now)
if err := c.PatchInteractiveCard(context.Background(), PatchCardParams{
InstallationID: testCreds(),
LarkCardMessageID: "om_patch_um",
CardJSON: render.JSON,
}); err != nil {
t.Fatalf("patch: %v", err)
}
assertCardContentHasUpdateMulti(t, capturedContent)
}
func assertCardContentHasUpdateMulti(t *testing.T, content string) {
t.Helper()
if content == "" {
t.Fatalf("captured content empty — fake server did not receive the request body")
}
var doc map[string]any
if err := json.Unmarshal([]byte(content), &doc); err != nil {
t.Fatalf("card content is not valid JSON: %v (raw=%s)", err, content)
}
cfg, ok := doc["config"].(map[string]any)
if !ok {
t.Fatalf("card content missing config block (raw=%s)", content)
}
if v, _ := cfg["update_multi"].(bool); !v {
t.Fatalf("config.update_multi must be true so the card is patchable on Lark's side; got config=%v (raw=%s)", cfg, content)
}
}
func TestHTTPClient_SendInteractiveCard_HappyPath(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_1", 7200)
fake.stubSend(
map[string]any{
"code": 0,
"msg": "ok",
"data": map[string]string{"message_id": "om_msg_42"},
},
func(r *http.Request, body map[string]string) {
if got := r.URL.Query().Get("receive_id_type"); got != "chat_id" {
t.Errorf("receive_id_type: got %q want chat_id", got)
}
if body["receive_id"] != "oc_chat_1" {
t.Errorf("receive_id: got %q", body["receive_id"])
}
if body["msg_type"] != "interactive" {
t.Errorf("msg_type: got %q want interactive", body["msg_type"])
}
if !strings.Contains(body["content"], "\"tag\"") {
t.Errorf("content not a card body: %q", body["content"])
}
},
)
c := newTestClient(fake, time.Now)
msgID, err := c.SendInteractiveCard(context.Background(), SendCardParams{
InstallationID: testCreds(),
ChatID: ChatID("oc_chat_1"),
CardJSON: `{"tag":"div","text":"hi"}`,
})
if err != nil {
t.Fatalf("send: %v", err)
}
if msgID != "om_msg_42" {
t.Errorf("message id: got %q want om_msg_42", msgID)
}
if got := fake.lastAuth(); got != "Bearer tok_1" {
t.Errorf("Authorization header: got %q want Bearer tok_1", got)
}
}
// TestHTTPClient_SendTextMessage_HappyPath pins the wire shape of the
// plain text outbound used for chat replies + /issue confirmations.
// Path, query, bearer auth, msg_type, and the double-JSON-encoded
// `content` envelope all matter — Lark rejects anything off-spec and
// the failures are silent-but-non-2xx, which is hard to debug
// in production without this kind of contract pin.
func TestHTTPClient_SendTextMessage_HappyPath(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_text", 7200)
fake.stubSend(
map[string]any{
"code": 0,
"msg": "ok",
"data": map[string]string{"message_id": "om_text_1"},
},
func(r *http.Request, body map[string]string) {
if r.URL.Path != "/open-apis/im/v1/messages" {
t.Errorf("path: got %q want /open-apis/im/v1/messages", r.URL.Path)
}
if got := r.URL.Query().Get("receive_id_type"); got != "chat_id" {
t.Errorf("receive_id_type: got %q want chat_id", got)
}
if body["receive_id"] != "oc_chat_42" {
t.Errorf("receive_id: got %q want oc_chat_42", body["receive_id"])
}
if body["msg_type"] != "text" {
t.Errorf("msg_type: got %q want text (NOT interactive — chat replies are plain bubbles)", body["msg_type"])
}
// content is a JSON-encoded string Lark requires: the outer
// HTTP body is JSON, and `content` is another JSON
// document INSIDE it. Decode and inspect.
var inner map[string]string
if err := json.Unmarshal([]byte(body["content"]), &inner); err != nil {
t.Fatalf("content is not valid inner JSON: %v (raw=%q)", err, body["content"])
}
if inner["text"] != "Hello world" {
t.Errorf("inner content.text: got %q want Hello world", inner["text"])
}
},
)
c := newTestClient(fake, time.Now)
msgID, err := c.SendTextMessage(context.Background(), SendTextParams{
InstallationID: testCreds(),
ChatID: ChatID("oc_chat_42"),
Text: "Hello world",
})
if err != nil {
t.Fatalf("send: %v", err)
}
if msgID != "om_text_1" {
t.Errorf("message id: got %q want om_text_1", msgID)
}
if got := fake.lastAuth(); got != "Bearer tok_text" {
t.Errorf("Authorization header: got %q want Bearer tok_text", got)
}
}
// TestHTTPClient_SendTextMessage_ReplyInThread pins the wire shape of a
// threaded reply: when ReplyTarget is set the client must POST to the
// reply endpoint (/messages/<id>/reply), carry reply_in_thread=true, and
// NOT include a chat-level receive_id — that's what lands the agent's
// reply inside the originating 话题 (thread) instead of the group.
func TestHTTPClient_SendTextMessage_ReplyInThread(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_reply", 7200)
fake.stubReply(
map[string]any{
"code": 0,
"msg": "ok",
"data": map[string]string{"message_id": "om_reply_1"},
},
func(r *http.Request, id string, body map[string]any) {
if id != "om_trigger" {
t.Errorf("reply target id: got %q want om_trigger", id)
}
if body["msg_type"] != "text" {
t.Errorf("msg_type: got %v want text", body["msg_type"])
}
if v, _ := body["reply_in_thread"].(bool); !v {
t.Errorf("reply_in_thread: got %v want true", body["reply_in_thread"])
}
if _, hasRecv := body["receive_id"]; hasRecv {
t.Errorf("reply endpoint body must NOT carry receive_id; got %v", body)
}
content, ok := body["content"].(string)
if !ok {
t.Fatalf("content missing or not a string: %v", body["content"])
}
var inner map[string]string
if err := json.Unmarshal([]byte(content), &inner); err != nil {
t.Fatalf("content inner JSON: %v (raw=%q)", err, content)
}
if inner["text"] != "threaded hi" {
t.Errorf("inner content.text: got %q want threaded hi", inner["text"])
}
},
)
c := newTestClient(fake, time.Now)
msgID, err := c.SendTextMessage(context.Background(), SendTextParams{
InstallationID: testCreds(),
ChatID: ChatID("oc_chat_42"),
Text: "threaded hi",
ReplyTarget: ReplyTarget{MessageID: "om_trigger", InThread: true},
})
if err != nil {
t.Fatalf("send: %v", err)
}
if msgID != "om_reply_1" {
t.Errorf("message id: got %q want om_reply_1", msgID)
}
}
// TestHTTPClient_SendMarkdownCard_HappyPath pins the wire shape of the
// schema-2.0 card we send for markdown chat replies. The MUST-haves:
// msg_type=interactive (not text), content is a JSON-encoded card
// envelope, the card has `schema: "2.0"` at the top level, and the
// body element is `{tag: "markdown", content: <agent's md verbatim>}`.
// Lark rejects malformed cards with a generic 9499xxxx code that's
// painful to root-cause in production, so we contract-pin every level.
func TestHTTPClient_SendMarkdownCard_HappyPath(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_md", 7200)
fake.stubSend(
map[string]any{
"code": 0,
"msg": "ok",
"data": map[string]string{"message_id": "om_md_1"},
},
func(r *http.Request, body map[string]string) {
if r.URL.Path != "/open-apis/im/v1/messages" {
t.Errorf("path: got %q want /open-apis/im/v1/messages", r.URL.Path)
}
if got := r.URL.Query().Get("receive_id_type"); got != "chat_id" {
t.Errorf("receive_id_type: got %q want chat_id", got)
}
if body["msg_type"] != "interactive" {
t.Errorf("msg_type: got %q want interactive (markdown cards ride the interactive endpoint)", body["msg_type"])
}
var card map[string]any
if err := json.Unmarshal([]byte(body["content"]), &card); err != nil {
t.Fatalf("content is not valid card JSON: %v (raw=%q)", err, body["content"])
}
if card["schema"] != "2.0" {
t.Errorf("card.schema: got %v want \"2.0\"", card["schema"])
}
bodyDoc, _ := card["body"].(map[string]any)
elements, _ := bodyDoc["elements"].([]any)
if len(elements) != 1 {
t.Fatalf("expected exactly one body element; got %d", len(elements))
}
el, _ := elements[0].(map[string]any)
if el["tag"] != "markdown" {
t.Errorf("element.tag: got %v want \"markdown\"", el["tag"])
}
if el["content"] != "# Heading\n- list" {
t.Errorf("markdown body must be forwarded verbatim; got %q", el["content"])
}
},
)
c := newTestClient(fake, time.Now)
msgID, err := c.SendMarkdownCard(context.Background(), SendMarkdownCardParams{
InstallationID: testCreds(),
ChatID: ChatID("oc_chat_42"),
Markdown: "# Heading\n- list",
})
if err != nil {
t.Fatalf("send markdown card: %v", err)
}
if msgID != "om_md_1" {
t.Errorf("message id: got %q want om_md_1", msgID)
}
if got := fake.lastAuth(); got != "Bearer tok_md" {
t.Errorf("Authorization header: got %q want Bearer tok_md", got)
}
}
// TestHTTPClient_SendTextMessage_EncodesSpecialCharacters guards the
// inner JSON envelope's escaping. Lark's spec is "content MUST be a
// JSON-encoded string", which means newlines and quotes have to be
// double-escaped — once when we marshal the inner `{"text": ...}`,
// then once more implicitly when the outer body is encoded for the
// HTTP request. Forgetting either pass corrupts the text Lark renders
// (or worse, rejects the message with a body parse error).
func TestHTTPClient_SendTextMessage_EncodesSpecialCharacters(t *testing.T) {
cases := []struct {
name string
text string
}{
{"multiline", "first line\nsecond line"},
{"double_quote", `she said "hi"`},
{"backslash", `path\to\file`},
{"chinese", "你好,世界 🌏"},
{"tab_and_newline", "col1\tcol2\nrow2"},
{"json_lookalike", `{"fake": "json"}`},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok", 7200)
fake.stubSend(
map[string]any{"code": 0, "data": map[string]string{"message_id": "om_x"}},
func(r *http.Request, body map[string]string) {
var inner map[string]string
if err := json.Unmarshal([]byte(body["content"]), &inner); err != nil {
t.Fatalf("content envelope not valid JSON after wire-encode round trip: %v (raw=%q)", err, body["content"])
}
if inner["text"] != tc.text {
t.Errorf("text round-trip failed\n got: %q\n want: %q", inner["text"], tc.text)
}
},
)
c := newTestClient(fake, time.Now)
if _, err := c.SendTextMessage(context.Background(), SendTextParams{
InstallationID: testCreds(),
ChatID: ChatID("oc_chat_1"),
Text: tc.text,
}); err != nil {
t.Fatalf("send: %v", err)
}
})
}
}
// TestHTTPClient_SendTextMessage_LarkErrorCode pins the failure path:
// non-zero `code` becomes a wrapped error; a missing message_id even
// with code=0 is still treated as failure (matches the success-card
// path so callers don't have to special-case the response shapes).
func TestHTTPClient_SendTextMessage_LarkErrorCode(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok", 7200)
fake.stubSend(
map[string]any{
"code": 234567,
"msg": "Permission denied",
},
nil,
)
c := newTestClient(fake, time.Now)
_, err := c.SendTextMessage(context.Background(), SendTextParams{
InstallationID: testCreds(),
ChatID: ChatID("oc"),
Text: "hi",
})
if err == nil {
t.Fatal("expected error on non-zero Lark code")
}
if !strings.Contains(err.Error(), "234567") {
t.Errorf("error should surface the Lark code; got %v", err)
}
}
func TestHTTPClient_SendInteractiveCard_TokenCached(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_cached", 7200)
fake.stubSend(
map[string]any{
"code": 0,
"data": map[string]string{"message_id": "om_msg_x"},
},
nil,
)
c := newTestClient(fake, time.Now)
for i := 0; i < 3; i++ {
if _, err := c.SendInteractiveCard(context.Background(), SendCardParams{
InstallationID: testCreds(),
ChatID: ChatID("oc_chat_1"),
CardJSON: `{}`,
}); err != nil {
t.Fatalf("iter %d: %v", i, err)
}
}
if got := fake.tokenN.Load(); got != 1 {
t.Errorf("token endpoint hits: got %d want 1 (cached after first call)", got)
}
if got := fake.sendN.Load(); got != 3 {
t.Errorf("send endpoint hits: got %d want 3", got)
}
}
func TestHTTPClient_TokenRefreshAfterExpiry(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_refresh", 120) // 120s expire → 60s usable after safety margin
fake.stubSend(
map[string]any{
"code": 0,
"data": map[string]string{"message_id": "om"},
},
nil,
)
now := time.Unix(1_700_000_000, 0)
clock := &fakeClock{now: now}
c := NewHTTPAPIClient(HTTPClientConfig{BaseURL: fake.URL(), Now: clock.Now}).(*httpAPIClient)
// First call — fetches token.
if _, err := c.SendInteractiveCard(context.Background(), SendCardParams{
InstallationID: testCreds(),
ChatID: ChatID("oc"),
CardJSON: `{}`,
}); err != nil {
t.Fatalf("first send: %v", err)
}
if fake.tokenN.Load() != 1 {
t.Fatalf("first call should have fetched a token, got tokenN=%d", fake.tokenN.Load())
}
// Advance past the cached token's expiry (token expire 120s,
// safety margin 60s → cache valid for 60s of wall-clock).
clock.Advance(90 * time.Second)
if _, err := c.SendInteractiveCard(context.Background(), SendCardParams{
InstallationID: testCreds(),
ChatID: ChatID("oc"),
CardJSON: `{}`,
}); err != nil {
t.Fatalf("post-expiry send: %v", err)
}
if got := fake.tokenN.Load(); got != 2 {
t.Errorf("token endpoint hits after expiry: got %d want 2", got)
}
}
func TestHTTPClient_SendInteractiveCard_LarkErrorCode(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_e", 7200)
fake.stubSend(map[string]any{"code": 230001, "msg": "no permission"}, nil)
c := newTestClient(fake, time.Now)
_, err := c.SendInteractiveCard(context.Background(), SendCardParams{
InstallationID: testCreds(),
ChatID: ChatID("oc"),
CardJSON: `{}`,
})
if err == nil {
t.Fatal("want error on non-zero code")
}
if !strings.Contains(err.Error(), "code=230001") {
t.Errorf("error should surface code: %v", err)
}
}
// TestHTTPClient_SendMethods_ReturnTypedAPIError pins that the three
// send methods used for threaded replies surface a non-zero Lark code
// as a structured *APIError, so the outbound fallback can classify
// "topic cannot receive this reply" codes without string matching.
func TestHTTPClient_SendMethods_ReturnTypedAPIError(t *testing.T) {
cases := []struct {
name string
call func(c *httpAPIClient) error
}{
{"interactive", func(c *httpAPIClient) error {
_, err := c.SendInteractiveCard(context.Background(), SendCardParams{InstallationID: testCreds(), ChatID: "oc", CardJSON: `{}`})
return err
}},
{"text", func(c *httpAPIClient) error {
_, err := c.SendTextMessage(context.Background(), SendTextParams{InstallationID: testCreds(), ChatID: "oc", Text: "hi"})
return err
}},
{"markdown", func(c *httpAPIClient) error {
_, err := c.SendMarkdownCard(context.Background(), SendMarkdownCardParams{InstallationID: testCreds(), ChatID: "oc", Markdown: "**hi**"})
return err
}},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_e", 7200)
fake.stubSend(map[string]any{"code": 230071, "msg": "group does not support reply in thread"}, nil)
c := newTestClient(fake, time.Now)
err := tc.call(c)
var apiErr *APIError
if !errors.As(err, &apiErr) {
t.Fatalf("want *APIError, got %T (%v)", err, err)
}
if apiErr.Code != 230071 {
t.Errorf("APIError.Code = %d; want 230071", apiErr.Code)
}
if !isThreadReplyUnsupported(err) {
t.Errorf("230071 should classify as thread-reply-unsupported")
}
if !strings.Contains(err.Error(), "code=230071") {
t.Errorf("error string should preserve code=230071: %v", err)
}
})
}
}
// TestIsThreadReplyUnsupported_ExcludesAmbiguous guards that ambiguous
// and rate-limit failures are NOT treated as classified thread errors,
// so they never trigger a chat-level fallback.
func TestIsThreadReplyUnsupported_ExcludesAmbiguous(t *testing.T) {
if isThreadReplyUnsupported(errors.New("transport failure")) {
t.Error("plain transport error must not classify as thread-reply-unsupported")
}
if isThreadReplyUnsupported(&APIError{Code: 230020, Msg: "rate limit"}) {
t.Error("rate limit (230020) must not classify as thread-reply-unsupported")
}
if isThreadReplyUnsupported(&APIError{Code: 230049, Msg: "message is being sent"}) {
t.Error("ambiguous 'being sent' (230049) must not classify as thread-reply-unsupported")
}
if !isThreadReplyUnsupported(&APIError{Code: 230072, Msg: "aggregated"}) {
t.Error("aggregated message (230072) should classify as thread-reply-unsupported")
}
}
func TestHTTPClient_SendInteractiveCard_TokenExpired_InvalidatesCache(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_first", 7200)
// First send replies with expired-token. Second send (after the
// client should have dropped its cache) reaches the token
// endpoint again. We swap the send handler mid-test to model
// this without race conditions: send fails first, second call
// from the same fake gets the token-endpoint hit + a fresh send
// reply. To keep the test small we simply assert tokenN
// increments after the failing call when the caller retries.
var sendCalls atomic.Int32
fake.mux.HandleFunc("/open-apis/im/v1/messages", func(w http.ResponseWriter, r *http.Request) {
fake.sendN.Add(1)
n := sendCalls.Add(1)
if n == 1 {
writeJSON(w, map[string]any{"code": codeTokenExpired, "msg": "expired"})
return
}
writeJSON(w, map[string]any{"code": 0, "data": map[string]string{"message_id": "om_ok"}})
})
c := newTestClient(fake, time.Now)
_, err := c.SendInteractiveCard(context.Background(), SendCardParams{
InstallationID: testCreds(),
ChatID: ChatID("oc"),
CardJSON: `{}`,
})
if err == nil {
t.Fatal("first send must fail with token-expired")
}
if !strings.Contains(err.Error(), "code=99991663") {
t.Errorf("error should mention token-expired code: %v", err)
}
// Caller's retry — should re-fetch the token, then succeed.
msgID, err := c.SendInteractiveCard(context.Background(), SendCardParams{
InstallationID: testCreds(),
ChatID: ChatID("oc"),
CardJSON: `{}`,
})
if err != nil {
t.Fatalf("retry send: %v", err)
}
if msgID != "om_ok" {
t.Errorf("retry message id: got %q", msgID)
}
if got := fake.tokenN.Load(); got != 2 {
t.Errorf("token endpoint hits after invalidation: got %d want 2", got)
}
}
func TestHTTPClient_PatchInteractiveCard_HappyPath(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_p", 7200)
fake.stubPatch(
map[string]any{"code": 0, "msg": "ok"},
func(r *http.Request, id string, body map[string]string) {
if id != "om_msg_42" {
t.Errorf("patch id: got %q want om_msg_42", id)
}
if !strings.Contains(body["content"], "updated") {
t.Errorf("patch content: %q", body["content"])
}
},
)
c := newTestClient(fake, time.Now)
if err := c.PatchInteractiveCard(context.Background(), PatchCardParams{
InstallationID: testCreds(),
LarkCardMessageID: "om_msg_42",
CardJSON: `{"text":"updated"}`,
}); err != nil {
t.Fatalf("patch: %v", err)
}
}
func TestHTTPClient_PatchInteractiveCard_LarkErrorCode(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_p", 7200)
fake.stubPatch(map[string]any{"code": 230002, "msg": "card not found"}, nil)
c := newTestClient(fake, time.Now)
err := c.PatchInteractiveCard(context.Background(), PatchCardParams{
InstallationID: testCreds(),
LarkCardMessageID: "om_msg_x",
CardJSON: `{}`,
})
if err == nil || !strings.Contains(err.Error(), "code=230002") {
t.Errorf("want code=230002 in error, got %v", err)
}
}
func TestHTTPClient_SendBindingPromptCard_HappyPath(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_b", 7200)
var capturedBody map[string]string
fake.mux.HandleFunc("/open-apis/im/v1/messages", func(w http.ResponseWriter, r *http.Request) {
fake.bindN.Add(1)
_ = json.NewDecoder(r.Body).Decode(&capturedBody)
if got := r.URL.Query().Get("receive_id_type"); got != "open_id" {
t.Errorf("receive_id_type: got %q want open_id", got)
}
writeJSON(w, map[string]any{"code": 0, "data": map[string]string{"message_id": "om_bind"}})
})
c := newTestClient(fake, time.Now)
if err := c.SendBindingPromptCard(context.Background(), BindingPromptParams{
InstallationID: testCreds(),
OpenID: OpenID("ou_user_1"),
BindURL: "https://multica.test/lark/bind?token=abc",
}); err != nil {
t.Fatalf("bind prompt: %v", err)
}
if capturedBody["receive_id"] != "ou_user_1" {
t.Errorf("receive_id: got %q", capturedBody["receive_id"])
}
if !strings.Contains(capturedBody["content"], "multica.test/lark/bind") {
t.Errorf("binding card should embed BindURL: %q", capturedBody["content"])
}
if !strings.Contains(capturedBody["content"], "去绑定") {
t.Errorf("binding card should carry the localized CTA: %q", capturedBody["content"])
}
}
func TestHTTPClient_TokenEndpointError(t *testing.T) {
fake := newLarkFake(t)
fake.stubTokenError(10003, "invalid app_id or app_secret")
c := newTestClient(fake, time.Now)
_, err := c.SendInteractiveCard(context.Background(), SendCardParams{
InstallationID: testCreds(),
ChatID: ChatID("oc"),
CardJSON: `{}`,
})
if err == nil || !strings.Contains(err.Error(), "code=10003") {
t.Errorf("want code=10003 surfaced, got %v", err)
}
}
func TestHTTPClient_AddMessageReaction_HappyPath(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_react", 7200)
fake.stubReaction(map[string]any{"code": 0, "msg": "ok", "data": map[string]string{"reaction_id": "re_42"}}, func(r *http.Request, id string, body map[string]any) {
if id != "om_user_msg_1" {
t.Errorf("message id: got %q want om_user_msg_1", id)
}
if got := r.Header.Get("Authorization"); got != "Bearer tok_react" {
t.Errorf("Authorization=%q want Bearer tok_react", got)
}
reactionType, ok := body["reaction_type"].(map[string]any)
if !ok {
t.Fatalf("reaction_type missing or wrong shape: %v", body)
}
if got := reactionType["emoji_type"]; got != "Typing" {
t.Errorf("emoji_type=%v want Typing", got)
}
})
c := newTestClient(fake, time.Now)
reactionID, err := c.AddMessageReaction(context.Background(), AddReactionParams{
InstallationID: testCreds(),
MessageID: "om_user_msg_1",
EmojiType: "Typing",
})
if err != nil {
t.Fatalf("AddMessageReaction: %v", err)
}
if reactionID != "re_42" {
t.Errorf("reaction id: got %q want re_42", reactionID)
}
if got := fake.reactN.Load(); got != 1 {
t.Fatalf("reaction endpoint calls=%d want 1", got)
}
}
func TestHTTPClient_DeleteMessageReaction_HappyPath(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_del", 7200)
fake.stubReactionDelete(map[string]any{"code": 0, "msg": "ok"}, func(r *http.Request, msgID string, reactionID string) {
if msgID != "om_user_msg_1" {
t.Errorf("message id: got %q want om_user_msg_1", msgID)
}
if reactionID != "re_42" {
t.Errorf("reaction id: got %q want re_42", reactionID)
}
if got := r.Header.Get("Authorization"); got != "Bearer tok_del" {
t.Errorf("Authorization=%q want Bearer tok_del", got)
}
})
c := newTestClient(fake, time.Now)
if err := c.DeleteMessageReaction(context.Background(), DeleteReactionParams{
InstallationID: testCreds(),
MessageID: "om_user_msg_1",
ReactionID: "re_42",
}); err != nil {
t.Fatalf("DeleteMessageReaction: %v", err)
}
if got := fake.delRN.Load(); got != 1 {
t.Fatalf("reaction delete endpoint calls=%d want 1", got)
}
}
func TestHTTPClient_AddMessageReaction_Validation(t *testing.T) {
c := NewHTTPAPIClient(HTTPClientConfig{}).(*httpAPIClient)
_, err := c.AddMessageReaction(context.Background(), AddReactionParams{MessageID: "m"})
if err == nil || !strings.Contains(err.Error(), "missing emoji_type") {
t.Errorf("want missing emoji_type error, got %v", err)
}
_, err = c.AddMessageReaction(context.Background(), AddReactionParams{EmojiType: "Typing"})
if err == nil || !strings.Contains(err.Error(), "missing message_id") {
t.Errorf("want missing message_id error, got %v", err)
}
}
func TestHTTPClient_DeleteMessageReaction_Validation(t *testing.T) {
c := NewHTTPAPIClient(HTTPClientConfig{}).(*httpAPIClient)
err := c.DeleteMessageReaction(context.Background(), DeleteReactionParams{ReactionID: "re"})
if err == nil || !strings.Contains(err.Error(), "missing message_id") {
t.Errorf("want missing message_id error, got %v", err)
}
err = c.DeleteMessageReaction(context.Background(), DeleteReactionParams{MessageID: "m"})
if err == nil || !strings.Contains(err.Error(), "missing reaction_id") {
t.Errorf("want missing reaction_id error, got %v", err)
}
}
func TestHTTPClient_MissingAppCredentials(t *testing.T) {
c := NewHTTPAPIClient(HTTPClientConfig{}).(*httpAPIClient)
_, err := c.tenantAccessToken(context.Background(), InstallationCredentials{AppSecret: "x"})
if err == nil || !strings.Contains(err.Error(), "app_id") {
t.Errorf("want missing app_id error, got %v", err)
}
_, err = c.tenantAccessToken(context.Background(), InstallationCredentials{AppID: "x"})
if err == nil || !strings.Contains(err.Error(), "app_secret") {
t.Errorf("want missing app_secret error, got %v", err)
}
}
func TestHTTPClient_MissingChatID_PreAuth(t *testing.T) {
// chat_id validation must short-circuit BEFORE any auth round-trip
// — otherwise a misuse leaks load to the token endpoint.
fake := newLarkFake(t)
c := newTestClient(fake, time.Now)
_, err := c.SendInteractiveCard(context.Background(), SendCardParams{
InstallationID: testCreds(),
CardJSON: `{}`,
})
if err == nil || !strings.Contains(err.Error(), "chat_id") {
t.Errorf("want missing chat_id error, got %v", err)
}
if got := fake.tokenN.Load(); got != 0 {
t.Errorf("token endpoint must not be hit on bad input: got %d", got)
}
}
func TestHTTPClient_MissingCardJSON(t *testing.T) {
c := NewHTTPAPIClient(HTTPClientConfig{}).(*httpAPIClient)
if _, err := c.SendInteractiveCard(context.Background(), SendCardParams{
InstallationID: testCreds(),
ChatID: ChatID("oc"),
}); err == nil || !strings.Contains(err.Error(), "card json") {
t.Errorf("send: want missing card json, got %v", err)
}
if err := c.PatchInteractiveCard(context.Background(), PatchCardParams{
InstallationID: testCreds(),
LarkCardMessageID: "om",
}); err == nil || !strings.Contains(err.Error(), "card json") {
t.Errorf("patch: want missing card json, got %v", err)
}
}
func TestHTTPClient_PatchMissingID(t *testing.T) {
c := NewHTTPAPIClient(HTTPClientConfig{}).(*httpAPIClient)
err := c.PatchInteractiveCard(context.Background(), PatchCardParams{
InstallationID: testCreds(),
CardJSON: `{}`,
})
if err == nil || !strings.Contains(err.Error(), "card message id") {
t.Errorf("want missing message id error, got %v", err)
}
}
func TestHTTPClient_BindingPromptValidation(t *testing.T) {
c := NewHTTPAPIClient(HTTPClientConfig{}).(*httpAPIClient)
if err := c.SendBindingPromptCard(context.Background(), BindingPromptParams{
InstallationID: testCreds(),
BindURL: "https://x",
}); err == nil || !strings.Contains(err.Error(), "open_id") {
t.Errorf("want missing open_id, got %v", err)
}
if err := c.SendBindingPromptCard(context.Background(), BindingPromptParams{
InstallationID: testCreds(),
OpenID: "ou",
}); err == nil || !strings.Contains(err.Error(), "bind url") {
t.Errorf("want missing bind url, got %v", err)
}
}
// TestHTTPClient_GetBotInfo_HappyPath drives the device-flow follow-up
// step: once RegistrationService has fresh client_id / client_secret
// from /oauth/v1/app/registration, it mints a tenant_access_token and
// asks /open-apis/bot/v3/info for the Bot's per-installation open_id,
// then resolves the bot's union_id via /open-apis/contact/v3/users/
// {open_id}?user_id_type=open_id. Both identifiers are persisted on
// the installation row; the union_id is what the WS decoder uses to
// route inbound @-mentions in multi-bot group chats (MUL-2671). The
// other fields on the bot/v3/info response (display name, avatar,
// IP whitelist) are deliberately dropped on the floor.
func TestHTTPClient_GetBotInfo_HappyPath(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_bi", 7200)
fake.mux.HandleFunc("/open-apis/bot/v3/info", func(w http.ResponseWriter, r *http.Request) {
if r.Method != http.MethodGet {
t.Errorf("bot info: want GET, got %s", r.Method)
}
if got := r.Header.Get("Authorization"); got != "Bearer tok_bi" {
t.Errorf("bot info: Authorization=%q want Bearer tok_bi", got)
}
writeJSON(w, map[string]any{
"code": 0,
"msg": "ok",
"bot": map[string]any{
"open_id": "ou_bot_42",
"app_name": "PersonalAgent",
"avatar_url": "https://example/avatar.png",
},
})
})
fake.mux.HandleFunc("/open-apis/contact/v3/users/ou_bot_42", func(w http.ResponseWriter, r *http.Request) {
if r.Method != http.MethodGet {
t.Errorf("contact users: want GET, got %s", r.Method)
}
if got := r.URL.Query().Get("user_id_type"); got != "open_id" {
t.Errorf("contact users: user_id_type=%q want open_id", got)
}
if got := r.Header.Get("Authorization"); got != "Bearer tok_bi" {
t.Errorf("contact users: Authorization=%q want Bearer tok_bi", got)
}
writeJSON(w, map[string]any{
"code": 0,
"msg": "ok",
"data": map[string]any{
"user": map[string]any{
"union_id": "on_bot_42_stable",
},
},
})
})
c := NewHTTPAPIClient(HTTPClientConfig{BaseURL: fake.URL()})
info, err := c.GetBotInfo(context.Background(), testCreds())
if err != nil {
t.Fatalf("GetBotInfo: %v", err)
}
if string(info.OpenID) != "ou_bot_42" {
t.Errorf("OpenID: got %q want ou_bot_42", info.OpenID)
}
if info.UnionID != "on_bot_42_stable" {
t.Errorf("UnionID: got %q want on_bot_42_stable", info.UnionID)
}
}
// TestHTTPClient_GetBotInfo_UnionIDLookupSoftFails covers the case
// where /contact/v3/users returns a non-zero code (e.g. the app's
// contact scope was never approved). The install must still succeed
// with an empty UnionID so the operator can backfill later instead
// of the QR flow failing outright. The decoder transitional fallback
// keeps single-bot installs working in the gap.
func TestHTTPClient_GetBotInfo_UnionIDLookupSoftFails(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_bi_softfail", 7200)
fake.mux.HandleFunc("/open-apis/bot/v3/info", func(w http.ResponseWriter, r *http.Request) {
writeJSON(w, map[string]any{
"code": 0,
"msg": "ok",
"bot": map[string]any{"open_id": "ou_bot_softfail"},
})
})
fake.mux.HandleFunc("/open-apis/contact/v3/users/ou_bot_softfail", func(w http.ResponseWriter, r *http.Request) {
writeJSON(w, map[string]any{"code": 99991002, "msg": "no permission"})
})
c := NewHTTPAPIClient(HTTPClientConfig{BaseURL: fake.URL()})
info, err := c.GetBotInfo(context.Background(), testCreds())
if err != nil {
t.Fatalf("GetBotInfo unexpectedly errored on soft-fail: %v", err)
}
if string(info.OpenID) != "ou_bot_softfail" {
t.Errorf("OpenID: got %q want ou_bot_softfail", info.OpenID)
}
if info.UnionID != "" {
t.Errorf("UnionID: got %q want empty (soft-fail leaves backfill to operator)", info.UnionID)
}
}
// TestHTTPClient_GetBotInfo_LarkErrorCode surfaces a non-zero Lark
// error code (e.g. 230003 = bot disabled) as a wrapped error so
// RegistrationService can fail the install cleanly instead of
// recording a row with bot_open_id="".
func TestHTTPClient_GetBotInfo_LarkErrorCode(t *testing.T) {
fake := newLarkFake(t)
fake.stubToken("tok_bi_err", 7200)
fake.mux.HandleFunc("/open-apis/bot/v3/info", func(w http.ResponseWriter, r *http.Request) {
writeJSON(w, map[string]any{"code": 230003, "msg": "bot disabled"})
})
c := NewHTTPAPIClient(HTTPClientConfig{BaseURL: fake.URL()})
_, err := c.GetBotInfo(context.Background(), testCreds())
if err == nil || !strings.Contains(err.Error(), "code=230003") {
t.Errorf("want code=230003 surfaced, got %v", err)
}
}
// TestHTTPClient_GetBotInfo_MissingCredentials short-circuits before
// any HTTP round-trip when the caller hands in zero-value credentials.
// A misuse here should NOT leak load to Lark's token endpoint.
func TestHTTPClient_GetBotInfo_MissingCredentials(t *testing.T) {
fake := newLarkFake(t)
c := newTestClient(fake, time.Now)
if _, err := c.GetBotInfo(context.Background(), InstallationCredentials{}); err == nil ||
!strings.Contains(err.Error(), "missing app credentials") {
t.Errorf("want missing credentials error, got %v", err)
}
if got := fake.tokenN.Load(); got != 0 {
t.Errorf("token endpoint must not be hit on bad input: got %d", got)
}
}
func TestHTTPClient_BadHTTPStatus(t *testing.T) {
fake := newLarkFake(t)
// Token returns success.
fake.stubToken("tok", 7200)
// Send replies with 500 + body — exercise the non-2xx branch.
fake.mux.HandleFunc("/open-apis/im/v1/messages", func(w http.ResponseWriter, r *http.Request) {
fake.sendN.Add(1)
w.WriteHeader(500)
_, _ = io.WriteString(w, "boom")
})
c := newTestClient(fake, time.Now)
_, err := c.SendInteractiveCard(context.Background(), SendCardParams{
InstallationID: testCreds(),
ChatID: ChatID("oc"),
CardJSON: `{}`,
})
if err == nil || !strings.Contains(err.Error(), "http 500") {
t.Errorf("want http 500 surfaced, got %v", err)
}
}
func TestHTTPClient_TokenExpire_ClampedToSafety(t *testing.T) {
// Lark returns expire=10s — well under the safety margin. The
// client must NOT cache a token that is already past its safe
// window; instead it clamps to 2× safety margin so the cached
// entry is at least usable for one safety margin of wall-clock.
fake := newLarkFake(t)
fake.stubToken("tok_short", 10)
fake.stubSend(map[string]any{"code": 0, "data": map[string]string{"message_id": "om"}}, nil)
now := time.Unix(1_700_000_000, 0)
clock := &fakeClock{now: now}
c := NewHTTPAPIClient(HTTPClientConfig{BaseURL: fake.URL(), Now: clock.Now}).(*httpAPIClient)
if _, err := c.SendInteractiveCard(context.Background(), SendCardParams{
InstallationID: testCreds(),
ChatID: ChatID("oc"),
CardJSON: `{}`,
}); err != nil {
t.Fatalf("send: %v", err)
}
clock.Advance(30 * time.Second) // still within clamped window
if _, err := c.SendInteractiveCard(context.Background(), SendCardParams{
InstallationID: testCreds(),
ChatID: ChatID("oc"),
CardJSON: `{}`,
}); err != nil {
t.Fatalf("send2: %v", err)
}
if got := fake.tokenN.Load(); got != 1 {
t.Errorf("token endpoint hits within clamped window: got %d want 1", got)
}
}
func TestBindingPromptTemplate_Shape(t *testing.T) {
raw, err := bindingPromptTemplate("https://multica.test/bind?token=abc")
if err != nil {
t.Fatalf("template: %v", err)
}
var doc map[string]any
if err := json.Unmarshal([]byte(raw), &doc); err != nil {
t.Fatalf("template json: %v", err)
}
// Shape check — top-level keys exist and elements is non-empty.
if _, ok := doc["config"]; !ok {
t.Errorf("missing config")
}
if _, ok := doc["header"]; !ok {
t.Errorf("missing header")
}
elements, ok := doc["elements"].([]any)
if !ok || len(elements) < 2 {
t.Fatalf("elements: want >=2, got %v", doc["elements"])
}
// Last element should be the action button carrying the URL.
last, _ := elements[len(elements)-1].(map[string]any)
if last["tag"] != "action" {
t.Errorf("last element should be action: %v", last)
}
actions, _ := last["actions"].([]any)
if len(actions) == 0 {
t.Fatalf("no actions in card")
}
btn, _ := actions[0].(map[string]any)
if btn["url"] != "https://multica.test/bind?token=abc" {
t.Errorf("button url: got %v", btn["url"])
}
}
// fakeClock is a minimal monotonic clock for tests that need to drive
// the cache TTL deterministically.
type fakeClock struct{ now time.Time }
func (c *fakeClock) Now() time.Time { return c.now }
func (c *fakeClock) Advance(d time.Duration) { c.now = c.now.Add(d) }