Java Strings Deep Dive: Interning, StringBuilder, Text Blocks, and the Concatenation Benchmark
What the string pool and intern() really do, what + compiles to on JDK 25 (invokedynamic, not StringBuilder), a JMH benchmark showing + is fine in one expression but quadratic in a loop, the four text-block rules, and what JEPs 430, 459 and 465 actually say about why string templates never shipped.
Two strings print the same text, == says false, and a code reviewer says “use equals” without explaining why the same comparison quietly works in the next test. Somebody else has a report builder that takes seconds to run and was ‘fixed’ by replacing + with StringBuilder everywhere, including in places where it changes nothing. A third developer read that Java was getting STR."Hello, \{name}" and wrote it into a design doc. All three are questions about what String really does, and this article answers them with programs that were compiled and run on JDK 25: where string literals live and what intern() does, what the + operator compiles to, what it costs next to StringBuilder (measured), the rules of text blocks, and what actually happened to string templates — checked against the JEP pages themselves.
Versions.JDK 25.0.4.1+1 (Temurin, LTS), JMH 1.37, JUnit Jupiter 5.11.0, run on a 2-vCPU x86-64 virtual machine that was shared with other jobs while measuring, so absolute timings are noisy; the ordering and the orders of magnitude are what to trust. All code and output below comes from the strings module of the java-core-examples repository. The JEP statuses quoted in the last sections were read from openjdk.org on 2026-10-01. For how a string is laid out in memory (Latin-1 versus UTF-16 bytes), see Compact Strings in Java 9; this article is about what you do with strings.
Equal text is not always the same object: the string pool
Java has two questions you can ask about two strings. a.equals(b) asks “do they contain the same characters?”. a == b asks “are they the same object?”. For strings the second answer depends on how each string was created, because the JVM keeps a table of canonical string instances — the string pool — and every string literal in your source is taken from it. Write "hello" twice, even in different classes, and you get one object. Write new String("hello") and you get a fresh copy outside the pool. The program in StringPoolDemo.java shows each case:
String a = "hello";
String b = "hello";
String c = new String("hello");
System.out.println("literal == literal : " + (a == b));
System.out.println("literal == new String : " + (a == c));
System.out.println("literal.equals(new String) : " + a.equals(c));
System.out.println("literal == c.intern() : " + (a == c.intern()));
Read the diagram left to right: a and b point at the single pooled object, c points at a separate copy with identical characters, which is why the transcript prints true for literal-to-literal, false for literal-to-new String, and true for equals. intern() is the bridge: it looks the string up in the pool and hands back the pooled instance (the last line of the transcript above). The habit to take away is simple — compare string contents with equals, never == — and the rest of this section is about the cases where == surprises people who think they understand it.
Constants are folded at compile time; runtime concatenation is not
If both operands of + are compile-time constants, javac computes the result itself and the result is a pooled literal. If either operand is only known at run time, the concatenation happens at run time and produces a new object. The demo builds "hel" + "lo" both ways (StringPoolDemo.java):
static final String CONST = "hel"; // compile-time constant variable
static final String NOT_CONST = new String("hel"); // not a constant
String folded = CONST + "lo"; // folded by javac into the constant "hello"
String runtime = NOT_CONST + "lo"; // built at run time
System.out.println("constant + literal == \"hello\" : " + (folded == a));
System.out.println("runtime concat == \"hello\" : " + (runtime == a));
System.out.println("runtime concat.intern() == \"hello\": " + (runtime.intern() == a));
The trap.static final String CONST = "hel" is a constant, so CONST + "lo" == "hello" is true. Replace the initialiser with new String("hel") and the field is still final and still a String, but NOT_CONST + "lo" == "hello" is false. Whether == works on a concatenation depends on whether every operand is a constant, which is invisible at the call site. The Java Language Specification defines which expressions qualify in section 15.29, Constant Expressions; this article only demonstrates the two cases above.
intern() returns the pooled instance for a string: if an equal string is already in the pool you get that one back, otherwise your string is added and you get it back. There are two surprises, and both are in the last two lines of the transcript (01-string-pool.txt):
String java = new StringBuilder("ja").append("va").toString();
System.out.println("built \"java\".intern() == the built instance : " + (java.intern() == java));
String unseen = new StringBuilder("never-seen-").append(System.nanoTime()).toString();
System.out.println("built unseen string .intern() == the built instance: " + (unseen.intern() == unseen));
built "java".intern() == the built instance : false
built unseen string .intern() == the built instance: true
The string "java" was already in the pool before main ran (something in the JDK’s start-up had interned it; I did not track down what), so interning our freshly built copy returned a different object and ours stayed outside the pool. A string nobody had ever interned was adopted as-is. Both results follow the definition, but the first one is why test code like assertSame(built.intern(), built) is fragile: it passes or fails depending on what else the JVM has already seen. The repository pins both behaviours in StringBehaviourTest (09-tests.txt).
Should you call intern() on your own data? This is opinion, backed by what the JVM reports about itself. The pool is a fixed-size hash table whose default size on this JVM is shown below; it is a real data structure with real lookup cost on every intern() call, and interning millions of distinct values (user IDs, request bodies) adds them all to it. Interning pays off when you hold a very large number of duplicates of a small set of strings for a long time and memory is your problem; the cheaper first step is usually to parse into an enum or a shared value object. If heap usage from duplicate strings is the actual symptom, the G1 collector has a separate deduplication feature, off by default here (07-jvm-flags.txt):
I did not benchmark intern() against a hand-written HashMap deduplicator or enable string deduplication, so this article makes no measured claim about which is faster. Measure with your own data before adopting either.
Going deeper
What the + operator compiles to: one invokedynamic call, not a StringBuilder
Old advice says “+ creates a StringBuilder behind your back”. That was true of the bytecode javac emitted up to Java 8. Since Java 9, javac emits a single invokedynamic instruction instead. The reason is in the summary of JEP 280, Indify String Concatenation (status Closed / Delivered, release 9): the change is meant to “enable future optimizations of String concatenation without requiring further changes to the bytecode emitted by javac”. The source we inspect is the same expression compiled twice (ConcatShapes.java):
public static String one(String name, int n) {
return "user=" + name + " n=" + n;
}
Compiled normally, and compiled with the hidden javac switch -XDstringConcat=inline that restores the pre-9 shape, javap -c shows this (04-javap-indy-vs-inline.txt):
The diagram is the whole point: with the default shape, the bytecode says only “join these pieces using this recipe”, and the JDK library decides at run time how to do it. The recipe is visible in the class file (05-bootstrap-method.txt), where \u0001 marks each operand slot:
Reference: the complete javap output for both shapes
Both listings are for method one(String, int) of ConcatShapes. Output file: 04-javap-indy-vs-inline.txt. JEP 280 names the indy and indyWithConstants javac flavours; the inline flavour used above is not documented there, I used it because it produced the old StringBuilder bytecode when I ran it, and I did not look up javac’s source to see whether it is a supported option.
One place where + is still the wrong tool. The invokedynamic call joins the pieces of one expression well. It cannot reach across loop iterations: s += x in a loop is still a fresh concatenation, producing a brand-new string each pass. That is the next section’s benchmark.
The concatenation benchmark: + is fine in one expression and quadratic in a loop
Two claims need measurement. First: does replacing a single + expression with an explicit StringBuilder make it faster? Second: what does += in a loop really cost? Before timing anything, the loop problem can be seen without a stopwatch. ConcatLoopDemo.java appends to a string in a loop and adds up how many characters the intermediate strings contain:
int big = 20_000;
long total = 0;
String t = "";
for (int i = 0; i < big; i++) { t += 'x'; total += t.length(); }
System.out.println("n=" + big + ": final length=" + t.length() + ", characters in all intermediate strings=" + total);
n=20000: final length=20000, characters in all intermediate strings=200010000
Twenty thousand iterations produced a 20,000-character string, yet the strings created along the way held about 200 million characters between them, because each pass copies everything so far into a new string. That is arithmetic, not an optimisation gone wrong: n passes copy roughly n²/2 characters in total. The JMH code for the first question is in ConcatBenchmark.java:
@Benchmark public String plusExpression() { return "user=" + name + " n=" + n + " ok"; }
@Benchmark public String explicitBuilder() {
return new StringBuilder().append("user=").append(name).append(" n=").append(n).append(" ok").toString();
}
@Benchmark public String stringFormat() { return String.format("user=%s n=%d ok", name, n); }
@Benchmark public String formattedMethod() { return "user=%s n=%d ok".formatted(name, n); }
The bars show the first answer: the plain + expression (20.4 ± 7.3 ns) and the hand-written StringBuilder (20.8 ± 2.6 ns) are statistically indistinguishable, so rewriting a single expression to use a builder buys nothing on JDK 25. Format-style methods are a different story — formatted() took about 164 ns and String.format about 286 ns, roughly eight and fourteen times the + expression here, presumably because they parse a format string on every call (I did not profile that). (String.format has a wide error margin, ± 134 ns, so treat its number as an order of magnitude.) The raw scores are in 08-jmh-raw.txt:
@Benchmark public String loopPlusEquals() {
String s = "";
for (int i = 0; i < loop; i++) s += "x";
return s;
}
@Benchmark public String loopBuilder() {
StringBuilder sb = new StringBuilder();
for (int i = 0; i < loop; i++) sb.append("x");
return sb.toString();
}
@Benchmark public String loopBuilderPresized() {
StringBuilder sb = new StringBuilder(loop);
for (int i = 0; i < loop; i++) sb.append("x");
return sb.toString();
}
The log scale hides how wide the gap is, so here are the numbers: at 100 iterations += took 2.0 µs against 0.3 µs for the builder; at 1,000 it took 73.8 µs against 2.3 µs; at 10,000 it took 8.9 ms against 24.5 µs — about 365 times slower. Growing the loop by 100× (100 to 10,000) grew the += time about 4,422× but the builder’s only about 91×, which is what quadratic versus linear looks like. The raw rows (08-jmh-raw.txt):
Do not over-read the presized builder.new StringBuilder(n) looks cheaper in every row, but the error margins overlap at 10,000 iterations (21 ± 17 µs presized against 24 ± 9 µs plain) and this machine was shared. The safe reading is that pre-sizing is at most a modest gain next to the difference between += and any builder. Noise in this benchmark is real: an earlier run on the same machine gave noticeably different absolute numbers with the same ordering.
Reference: how the JMH runs were set up, and the Latin-1 versus UTF-16 rowsConcatBenchmark uses 2 forks, LoopConcatBenchmark 1 fork; both run 3 warm-up and 5 measurement iterations of 1 second in average-time mode. JMH takes a lock at /tmp/jmh.lock and refuses to start while another JMH process holds it; on a shared machine, wait and retry rather than forcing it. The annotation-processor block in the module’s pom.xml is required on JDK 25, which stopped discovering processors implicitly.
The benchmark also joins three copies of a Latin-1 string and three copies of a string containing a euro sign, which forces UTF-16 storage (the layout is explained in the Compact Strings article). Result: 15.8 ± 5.0 ns for Latin-1 and 23.2 ± 4.3 ns for UTF-16. The wide string looks somewhat slower, but the intervals nearly touch, so treat it as a hint and not as a measured ratio.
Going deeper
Text blocks: a multi-line String whose margin is decided by the closing quotes
A text block is a string literal that starts with """ followed by a line break and ends with """. It was delivered as a final feature in Java 15 (JEP 378, Closed / Delivered, release 15), it produces an ordinary String, and it exists so that JSON, SQL and HTML do not need a backslash-n and a quote-escape on every line. Four rules explain nearly everything that surprises people, and TextBlockDemo.java demonstrates each. The demo prints strings with spaces shown as dots and newlines shown as \n, so whitespace is visible.
Rule 1: the compiler strips the indentation all lines share, and the closing delimiter counts
String a = """
{
"id": 1
}
""";
String b = """
{
"id": 1
}""";
String c = """
left
""";
--- closing delimiter on its own line, aligned with the text ---
{\n
.."id":.1\n
}\n
--- closing delimiter right after the last line (no final newline) ---
{\n
.."id":.1\n
}
--- closing delimiter 2 columns left of the text: 2 spaces of indentation survive ---
..left\n
The diagram mirrors the transcript. The JLS calls what gets removed incidental white space: the compiler finds the smallest indentation among all content lines and the line holding the closing delimiter, and strips that much from every line. That is why the first two blocks produce {, "id": 1, } without any leading indentation, differing only in the final newline, and why pulling the closing quotes two columns left (third block) makes two spaces survive on every line. The delimiter is a margin ruler as well as a terminator.
Rule 2: trailing spaces vanish unless you escape them; rule 3: escapes run after stripping
String d = """
x \s
y\040
""";
show("trailing spaces are stripped, but \\s and \\040 survive", d);
String e = """
one \
line
""";
show("backslash at end of line joins the lines", e);
String f = """
tab:\t|
""";
show("escapes are processed after indentation is stripped", f);
String g = """
say "hi" and ""\"
""";
show("double quotes need no escape; only a run of three does", g);
--- trailing spaces are stripped, but \s and \040 survive ---
x....\n
y.\n
--- backslash at end of line joins the lines ---
one.line\n
--- escapes are processed after indentation is stripped ---
tab: |\n
--- double quotes need no escape; only a run of three does ---
say."hi".and."""\n
Trailing spaces on a line are stripped, which is a blessing for editors that trim whitespace and a trap when the trailing space matters. \s (a space) and \040 (an octal escape) both survive, as the x.... and y. lines show. Because escapes are translated after stripping, an escaped space at the end of a line cannot be eaten by the stripping step. A backslash at the very end of a line joins it with the next, which is how you write a long single-line string across several source lines. Double quotes need no escaping unless three appear in a row.
Rule 4: it is an ordinary String
String same = """
{
"id": 1
}
""";
System.out.println("two identical text blocks are the same pooled instance: " + (a == same));
System.out.println("equals the hand-written literal: " + a.equals("{\n \"id\": 1\n}\n"));
String h = """
Hello, %s! You are %d.
""".formatted("Ann", 30);
two identical text blocks are the same pooled instance: true
equals the hand-written literal: true
--- formatted() is the built-in way to interpolate ---
Hello,.Ann!.You.are.30.\n
Two identical text blocks are the same pooled instance (connecting back to the first section), and a text block equals the hand-written literal. For interpolation, the built-in tool is formatted(), which is String.format with the receiver as the format string — and from the benchmark above you know it is not free. The demo’s claims are pinned by assertions in StringBehaviourTest.
Text blocks and regular expressions. Text blocks are convenient for multi-line patterns, but the escape rules above apply to them: a regex \d must still be written \\d, and the line-joining backslash means a pattern ending a line with \ does not do what a regex reader expects. The Java Regex Tutorial covers the regex side.
String templates: what JEPs 430, 459 and 465 actually say, and why you cannot use them
String templates were supposed to be Java’s answer to interpolation: STR."Hello, \{name}". If you try that on JDK 25 today, here is what happens (source: StringTemplate.java, output: 06-string-template-compile-error.txt):
Both attempts fail, with and without --enable-preview, and the message is the compiler no longer recognising the \{ escape at all. So this is not a feature that is merely still in preview on JDK 25; the syntax is gone from the compiler. I tested JDK 25 only. What do the primary sources say about how it got there? I read each JEP page on openjdk.org and quote what is there, without filling in gaps:
What the pages say, exactly. JEP 430 (String Templates, Preview): Status Closed / Delivered, Release 21. JEP 459 (Second Preview): Status Closed / Delivered, Release 22; its History section says the second preview was proposed “to gain additional experience and feedback” and that, “except for a technical change in the types of template expressions, there are no changes relative to the first preview”. JEP 465 (Third Preview): Status Closed / Withdrawn, with no release listed. The page body still says “We here propose to finalize the feature with no further changes” under a Third Preview title — the page was evidently not rewritten when the plan changed, and it does not state a reason for the withdrawal. I want to be clear about that: the JEP pages record that it was withdrawn, not why. The JDK 23 project page lists its final feature set, and no string-template JEP is on it.
The reasons come from the Amber mailing list. Gavin Bierman wrote on 5 April 2024 (message in the amber-spec-experts archive): “there is support for a change in the design but a lack of clear consensus on what that new design might look like, the prudent course of action is to (i) NOT ship the current design as a preview feature in JDK 23, and (ii) take our time continuing the design process”, and, to be unambiguous, “there will be no string template feature, even with –enable-preview, in JDK 23.” The same thread quotes Brian Goetz’s earlier note of 8 March 2024, which says that experience using the feature in the jextract project showed “the role of processors is ‘outsized’ to the value they offer, and, after further exploration, we now believe it is possible to achieve the goals of the feature without an explicit ‘processor’ abstraction at all”. In other words: the processor concept (STR, FMT, RAW) was judged too heavy, and a redesign without it was being explored. Whether a new design has since been proposed, I did not research; check the JEP index before writing about it.
What this means for your code. Do not write STR."..." in anything: it does not compile on JDK 25, and on JDK 23 and later the preview was not shipped. If you read an older blog post that uses it, that post describes a JDK 21 or 22 preview that has been removed, not a feature waiting to be finalised. Today’s options are the ones already covered: + for short expressions, StringBuilder for loops, formatted() or String.format when you want a template-like format string and can afford the parse cost (measured above), and MessageFormat for locale-aware messages (not benchmarked here).
Honest advice, with the evidence behind each line. (1) Compare strings with equals; == happens to work for literals and constants and breaks for anything built at run time. (2) Do not sprinkle intern(); it is a lookup in a finite table, and the "java" example shows its result depends on what the JVM already holds. (3) Write + freely inside one expression; on JDK 25 it measured the same as a hand-written StringBuilder. (4) Use a StringBuilder when you append in a loop; the benchmark shows the gap growing to a few hundred times at 10,000 iterations. (5) Use text blocks for multi-line literals and remember the closing quotes set the margin. (6) Do not plan around string templates. Items (3) and (4) were measured on a shared 2-vCPU machine; the conclusions rest on gaps that are far larger than the noise, but your exact nanoseconds will differ.
No Comments yet!