1
0
Files
spring-grpc-boot4/docs/01-troubleshooting.md
Ankur 8547981634 Spring gRPC on Spring Boot 4: four call types, exception mapping, and a symptom-first troubleshooting suite
Companion code for the ankurm.com guide. Verified on Spring Boot 4.1.0, spring-grpc 1.1.0,
grpc-java 1.80.0, protobuf-java 4.34.2, JDK 25.0.3. results-full.txt is unedited mvn test
output: 11 tests, 0 failures. No installs needed - protoc and the gRPC codegen plugin resolve
as Maven artifacts, and every test uses the in-process transport.

_01_basics       all four call types and how their failure modes differ: iterator semantics on
                 server streaming, the half-close that client streaming hangs without, and the
                 independence of the two streams in bidi.
_02_troubleshooting
                 four failures reproduced then fixed - the 4 MB message limit and which side
                 enforces it, the absent default deadline, cancellation that never interrupts a
                 thread, and errors that arrive as UNKNOWN with no description.
_03_exceptions   @GrpcAdvice / @GrpcExceptionHandler: domain exception to NOT_FOUND with
                 trailers, validation to INVALID_ARGUMENT, and a catch-all that returns a
                 deliberate INTERNAL without leaking the original message across the boundary.

docs/01          symptom-first troubleshooting index, each entry marked [tested] or [documented]
docs/02          complete spring.grpc.* property reference, and the split between the Boot 4.1
                 integration (org.springframework.boot, version 4.1.0, owns properties and
                 auto-configuration) and the Spring gRPC project (org.springframework.grpc,
                 version 1.1.0, owns the programming model). The older standalone
                 spring-grpc-spring-boot-starter stops at 1.0.3 and is what most search results
                 describe.

Findings worth the commit message
  - The in-process transport CANNOT enforce message size limits: it passes messages by
    reference and never serialises them. A 4 MB + 1 KB message goes through cleanly in tests
    and fails in production with RESOURCE_EXHAUSTED. The test asserts this rather than
    pretending otherwise. Same blind spot covers compression, TLS, keepalive and LB.
  - Two property names cost real time while writing this:
    spring.grpc.server.inprocess.name (not in-process) and
    spring.grpc.client.channel.<n>.target (singular channel, and target not address).
    The second failure surfaces as UnknownHostException on the channel NAME.
  - gRPC still has no default deadline, and cancellation only sets a Context flag.
2026-07-31 23:05:02 +05:30

189 lines
8.4 KiB
Markdown

# Troubleshooting Spring gRPC on Boot 4
Symptom-first. Each entry is something whose error message points somewhere other than its cause.
Entries marked **[tested]** are reproduced in
[`HardToDiagnoseTest`](../src/test/java/com/ankurm/grpc/_02_troubleshooting/HardToDiagnoseTest.java);
entries marked **[documented]** need a real network and are stated rather than asserted, because a
test claiming to prove them over in-process transport would be lying.
---
## `UNAVAILABLE: Unable to resolve host <name>` — **[tested]**
**Cause, 90% of the time: a typo in a property name**, not a networking problem.
An unmatched channel name is passed straight to the DNS resolver as a target. Boot does not warn
that the channel is unconfigured. The exact names, both of which are easy to get wrong:
```properties
spring.grpc.server.inprocess.name=orders-test # NOT "in-process"
spring.grpc.client.channel.orders.target=static://... # SINGULAR "channel"; "target", NOT "address"
```
Also check `spring.grpc.client.inprocess.enabled=true` if you are using the in-process transport
from the client side.
## `IllegalStateException: No grpc channel factory found that supports target : <name>`
`GrpcChannelFactory` is a composite that asks each registered factory `supports(target)` **before**
named-channel indirection is applied. Passing a bare logical name to `createChannel()` can therefore
fail even though the name is correctly configured. In application code, inject the stub (via
`@ImportGrpcClients`) and let Boot resolve the target; in tests, pass the full target.
## `RESOURCE_EXHAUSTED: gRPC message exceeds maximum size 4194304` — **[tested]**
The default limit is **4 MB, per message, on the receiver**, and the two sides are configured
separately.
```properties
# large RESPONSE -> the CLIENT is the receiver
spring.grpc.client.channel.<name>.inbound.message.max-size=16MB
# large REQUEST -> the SERVER is the receiver
spring.grpc.server.inbound.message.max-size=16MB
```
Per call: `stub.withMaxInboundMessageSize(16 * 1024 * 1024)`.
Half of all "I raised the limit and it still fails" reports are the asymmetry: raising it on the
server does nothing for a large response.
> **The testing trap.** The in-process transport passes messages **by reference and never
> serialises them**, so it *cannot* enforce size limits. A payload-size regression passes every
> in-process test you have and fails in production. The same applies to compression and anything
> else depending on the wire format. This is demonstrated — the test asserts that a 4 MB + 1 KB
> message goes through in-process without complaint.
Prefer streaming to large unary messages: the limit is per message, so 1,000 small messages are
fine. Do not raise the limit globally — a large limit turns a malformed request into an OOM.
## Calls that hang forever — **[tested]**
**gRPC has no default deadline.** A call without one waits indefinitely. Nothing warns you, and in
testing the server always responds quickly, so it surfaces the day something downstream is slow —
as a thread pool that fills up with no errors in the log.
```java
stub.withDeadlineAfter(2, TimeUnit.SECONDS).getOrder(request);
```
Per channel: `spring.grpc.client.channel.<name>.default.deadline=2s`, or register a
`DefaultDeadlineSetupClientInterceptor`.
Deadlines are **absolute and propagate**: if A calls B with 2s remaining, B sees 2s, not a fresh 2s.
Set the deadline at the edge; do not re-set it at every hop, or a deep chain multiplies its budget.
## The server keeps working after the client gave up — **[tested]**
gRPC **does not interrupt your thread** on cancellation. It sets a flag on the `Context`.
```java
for (Item item : items) {
if (Context.current().isCancelled()) return; // no onCompleted/onError -- stream is closed
responseObserver.onNext(convert(item));
}
```
For blocking work handed to another thread, propagate with `Context.current().wrap(runnable)`
otherwise the flag is invisible there. Calling `onNext`/`onCompleted` after cancellation throws
`IllegalStateException`.
## Every error is `UNKNOWN` with no message — **[tested]**
Any exception escaping a handler becomes `UNKNOWN` with a **null description** — deliberately, since
leaking exception text across a service boundary is an information-disclosure risk.
```java
Metadata trailers = new Metadata();
trailers.put(REASON_KEY, "ORDER_LOCKED");
responseObserver.onError(new StatusRuntimeException(
Status.FAILED_PRECONDITION.withDescription("order is locked"), trailers));
```
Choose the code carefully — it is the contract that retry policies, circuit breakers and dashboards
key off:
| Do **not** retry | Retry may help |
|---|---|
| `INVALID_ARGUMENT`, `NOT_FOUND`, `ALREADY_EXISTS`, `FAILED_PRECONDITION`, `PERMISSION_DENIED`, `UNAUTHENTICATED` | `UNAVAILABLE`, `DEADLINE_EXCEEDED`, `RESOURCE_EXHAUSTED`, `ABORTED` |
Getting this wrong makes non-retryable failures look retryable and turns one bad request into a
storm.
## `IllegalStateException: call already closed` / `half-closed`
You called `onNext` or `onCompleted` after the stream ended. Usual causes:
- responding after cancellation or deadline expiry (check `Context.current().isCancelled()` first);
- calling `onCompleted()` twice;
- calling `onError()` and then `onCompleted()``onError` is terminal;
- two threads writing to one `StreamObserver`. **`StreamObserver` is not thread-safe.** Concurrent
`onNext` corrupts the stream, and the symptom is usually a deserialization error on the *far*
side, pointing nowhere near the bug.
## A client-streaming call never completes
You forgot the half-close. `requestObserver.onCompleted()` is what tells the server "no more
requests"; without it the server's `onCompleted` never fires and the call hangs until the deadline —
or forever, if you did not set one.
## Retries do nothing — **[documented]**
gRPC retries are configured through the **service config**, not a client API, and are off by
default:
```properties
spring.grpc.client.channel.orders.service-config.methodConfig[0].retryPolicy.maxAttempts=4
spring.grpc.client.channel.orders.service-config.methodConfig[0].retryPolicy.initialBackoff=0.5s
spring.grpc.client.channel.orders.service-config.methodConfig[0].retryPolicy.retryableStatusCodes[0]=UNAVAILABLE
```
Three reasons they silently do nothing:
1. The status you are failing with is not in `retryableStatusCodes`.
2. The response was already partially delivered — gRPC will not retry a committed stream.
3. Retry is disabled on the channel (`enableRetry()` / the corresponding property).
Only make idempotent methods retryable. gRPC has no idea whether your RPC is safe to repeat.
## Connections drop every N minutes through a proxy — **[documented]**
Load balancers and ingress controllers close idle HTTP/2 connections. gRPC reuses one long-lived
connection, so an idle stream is not an idle *connection* — the proxy disagrees.
```properties
spring.grpc.client.channel.<name>.keepalive.time=30s
spring.grpc.client.channel.<name>.keepalive.timeout=5s
spring.grpc.server.keepalive.connection.max-idle-time=...
spring.grpc.server.keepalive.connection.max-age=...
```
Keep the client's keepalive interval **shorter** than the proxy's idle timeout. Note the server can
reject keepalives that are too frequent (`ENHANCE_YOUR_CALM` / `too_many_pings`), which looks like a
random disconnect — so both ends need to agree.
## All traffic goes to one backend — **[documented]**
The default load-balancing policy is `pick_first`, not `round_robin`. With a DNS target resolving to
several addresses, a gRPC client picks one connection and keeps it. This is correct behaviour and
almost never what people expect behind a headless Kubernetes service.
Set `round_robin` via service config, and remember that DNS re-resolution happens on connection
failure, not on a timer — so scaling up does not redistribute existing connections.
---
## Test transport limitations, collected
The in-process transport is excellent for testing protocol behaviour and misleading for anything
wire-related. It **does** exercise interceptors, metadata, status codes, deadlines, cancellation and
flow control. It does **not** exercise:
- message size limits (no serialization — proven in the test suite),
- compression,
- TLS and credentials negotiation,
- keepalive, GOAWAY, idle timeouts,
- name resolution and load balancing,
- HTTP/2 framing and header limits.
Anything in that second list needs a real port, and belongs in a smaller, slower test tier.