WRITING · JOURNAL
Tianli Zeng · Writing
36 POSTSEngineering notes, water resources, and scattered thoughts. Also the hub for everything I write.

Turning Legacy Code and a Cabinet of Documents into a Domain Agent That Actually Works: The Six-Step Path
You have nine domain algorithms that have run for years and a pile of docx/pdf documents. How do you turn them into an agent that actually does your work? A tutorial you can follow: six steps, each covering just three things — what to do, how to know it worked, and where the trap is. The most expensive lesson: your verifiers lie too — three of four red acceptance runs were the gate's fault, not the answer's.

Zero to TestFlight: The Complete Path to Shipping an iPhone App Solo
A tutorial you can actually follow: five steps from account registration to TestFlight, and for each step just three things — what to do, how to know it worked, and where the trap is. The most expensive lesson happens before step one: an irreversible decision that burned two of my app identifiers.

Why They Have to Be Right
You lost ten pounds; she says it's the camera angle. She isn't judging you — she's doing bookkeeping. Read the ledger and you'll waste a lot less anger.

I Built My Son an 'Asset Dashboard': 83 House Rules Turned into a Ledger That Goes Up and Down
The problem with stickers and gold stars isn't that they don't work — it's that they don't accumulate. I turned our family's 83-rule points chart into a stock-market-style ledger: a trend line, an all-time-high badge, a rewards shop, and an auditable transaction log. This post shows what it looks like, which loopholes it closes, plus half a century of token-economy research and what I borrowed from Duolingo and ClassDojo.

我给三年级的孩子做了个练习站:题目无限,但每一套都有头有尾
一份期末卷长出来的练习站。一套 12 题、题号条、错题本记题型不记题目、经验值和勋章墙、每道题都能「再造一道类似的」。这篇把它有什么功能、长什么样、怎么用,一次讲清楚,配图全是真页面截的。

I Built My Kid His Own Market Dashboard, and He Checks It Every Day to See How Much He's Up
The problem with stickers and gold stars isn't that they don't work, it's that they don't add up — by the seventeenth one, nobody remembers what the first one was for. So I turned it into a ledger: every point earned or docked has a reason, a timestamp, and the balance at that moment, with the running total sitting up top, like watching a ticker.

I Can Generate Fifty Thousand Questions. My Kid Asked How Many There Are.
One third-grade final exam, and a practice site I built for my kid. Four teardowns in five days: infinite stream, replaying wrong answers, mistaking generators for a question bank, and a combinatorial explosion. Every teardown had the same root cause—I kept treating 'technically possible' as 'this is what the product should be.' Here's the whole build, and how to use it.

I Invented Six Numbers, and Every Gate Was Green
Turning a paper rule sheet into a config file, the most dangerous move isn't getting the math wrong. It's filling in a plausible-looking number where the original gave none. Bad math gets caught by tests; invented numbers don't — they look exactly like the real ones.

Moving the AI Into Excel: A Workflow Migration Forced by 'Regenerate'
After a script-generated workbook flattened my hand-tuned charts twice, I moved the AI from generating files outside to working in place inside Excel's sidebar — and the hard part wasn't the tool, it was repackaging two years of hard-earned rules for an assistant with no memory.

My Checker Gave 102 Green Ticks. You Opened the File to a Screen Full of Errors
I shipped the file with a note: machine-checked, 102 values verified. They opened it and every cell was broken. My checker was not lying — it was answering a different question than I thought.

I Drew a Map of My Own AI Setup — The TL Harness Manual
228 parts, 677 notes, and I couldn't tell you which ones were still turning. An inventory tells you what you have; it can't answer "what breaks if I delete this?" So I drew the relationships — and found that 419 of my notes were mislabeled while every single count added up perfectly.

I Keep 35 Alarms Watching Me. I Silenced the Most Important One Myself.
To make one alarm less noisy, I added a line that skipped repeats. From that moment it went permanently silent — for every conversation on this machine — and it still looked like it was running fine.

It Printed Output Every Single Run — That's Why Nobody Noticed It Was Missing 86%
A script ran for two months. Every run produced output. Every output looked fine. It was missing 86% of my text messages — and I had been using that output to make decisions.

过度乐观,就是悲观
你感受到的从来不是现实本身,是现实减去预期。现实好过预期才有多巴胺,不如预期就只有失望。所以过度乐观就是悲观——预期抬到天上,现实天天让你失望。教条式的乐观和自信,本质是把预期拉满,在批量给自己制造失望。而两千年前的斯多葛给出了另一种乐观:预期压到最坏,接受一切并从中学习——高预期容易 miss,低预期容易 beat。

情绪价值是个强化信号
情绪价值不是有用没用的问题,它是个强化信号——本身没有方向。绑在真实能力上,它是燃料;和能力脱钩,它是高能耗、不稳定、还会把你锁在错误的行为上。而当前这个内卷、下行、上升通道变窄的环境,正在系统性地让它脱钩——所以它的负作用在放大。

靠结构,不靠记性
重复的错误不该靠记性解决,该靠结构解决。教训记在笔记里是愿望,写进拦截钩子里才是制度。三个亲身的坑——一个看不见的字体、一次并发提交、一份报错的台账——机制是同一个:只要'做对'还依赖某个人当场不出错,错误就永远有复发的地方。把这个依赖拆掉。

CRITIC Objective Weighting: What It Rewards, and When It Quietly Breaks
CRITIC claims to let the data set the weights — but what does it actually reward? A full worked example, plus two failure modes almost nobody mentions: with only 2 indicators it degenerates exactly to std-ratio weighting, and with too few samples the weights are decorative.

The Benchmark Bill for Three Subscriptions: K3 for Volume, Claude for Quality, GPT for Spikes
Kimi K3 shipped with 14 head-to-head benchmarks. I transcribed both charts line by line and mapped them against my own usage: K3's five wins are all high-frequency grunt work (long-horizon agentic, web research, automation), Fable 5's six wins are all desk-quality work, GPT-5.6's two wins are peak-brain work. Route by task personality: K3 for volume, Claude for quality, GPT for spikes — full routing table inside.

Three Open-Source Contributions, In Retrospect: Raycast · WeChatTweak · Cardinal
Three PRs I sent upstream in 2026 — a Raycast extension, a WeChat 4.x binary patch, and persistent search history for a Rust+Tauri app. Three ecosystems, three different bars for 'getting merged,' and in every case the code itself was the easy part.

Say It Clearly: A Full Postmortem of Working with AI
I co-wrote a technical proposal with AI — dozens of rounds from first draft to final. The potholes weren't the valuable part; the postmortem was: almost every one traced back to the same thing — when, and how clearly, the initial conditions and boundary conditions were stated. One sentence before kickoff beats ten rounds of corrections. Ends with a 30-second briefing checklist that works for AI and human collaborators alike.

Open-Sourcing Two Native macOS Apps: Ask Claude and HydroMac
Two native macOS apps from my personal fleet, now open source: Ask Claude (chat on your Claude subscription, zero API key) and HydroMac (SwiftUI shell + Rust compute core for water engineering). Plus the hardest lesson before going public — data sanitization is far more than find-and-replace.

A Guided Tour of This Site
tianli.cyou just went through a 'paper manual' redesign and a serious spring cleaning: dead navigation cut, zombie services removed, private entry collapsed into a single bookmark — a guided tour, plus how one person maintains a fleet of sites.

How Many Pressure Gauges Does It Take to Pin a Leak to One Pipe? A Full Walk-Through of the Pressure-Fingerprint Method
On a 6-node teaching network, worked from the Hazen-Williams equation all the way to the sensor-count decision: how a handful of pressure readings become a ranked list of suspect pipes, and how a marginal-benefit curve answers how many gauges and where.

陵阳公样智能辅助设计研究
目的 针对唐代陵阳公样装饰纹样在现代传承中面临的创新性不足与公众参与度低等问题,探索一种基于人工智能的智能辅助设计方法,以实现传统纹样的活态传承与创造性转化。 方法 首先,基于皮尔斯符号学三元理论,从表征、客体和诠释三个维度对陵阳公样进行系统解析,并运用分裂语法对纹样结构进行层级化解构,构建了包含6...

Prompt Caching 架构:为什么 CLAUDE.md 越稳定越省钱
--- Prompt Caching 是 Claude API 层面的优化机制,核心原理简单直接:缓存按前缀匹配工作。 每次请求发给 Claude 的完整输入大致是这样的结构: 如果两次请求的"前缀"相同——即从第一个 token 开始到某个位置完全一致——第二次请求可以复用缓存,跳过对这段前缀的重...

浙东引水工程受水区降雨趋势与多尺度变率分析
ZENG Tian-li$^1$, ZUO Xiao-xia$^1$, YANG Yu$^1$, DH$^1$, WU Mu-hong$^2$, ZHONG Lü-bin$^2$, CHEN Shu-yang$^3$ (1. Zhejiang Design Institute of Water Co...

Claude Code 架构 · Layer 1:长期上下文层(CLAUDE.md / Memory)
长期上下文=每次对话自动加载的背景:CLAUDE.md 全局/项目规则 + Memory 跨会话记忆。它是 Claude Code 的地基,写好了 Agent 在任何会话都知道技术栈、规矩、偏好。

A Methodology for Water Resources Carrying Capacity Evaluation: AHP-CRITIC-TOPSIS in Practice
From indicator system design to combined weighting and final ranking — a reproducible technical pipeline for water resources carrying capacity evaluation, with Python implementation notes.

Building an Enterprise-Grade Claude Code Harness
From 43 commands to 14 skills, from hook automation to MCP integration — a complete engineering methodology for LLM toolchains.

MCP in Practice: An Agent Ecosystem from Code Search to Gmail
Three MCP servers, 512+ session records, semantic search — how to build an AI Agent toolchain that actually gets used daily.

Cost Engineering: Running 28 Production Services for $31/Month
One $30/month VPS plus a $1/year domain, hosting 24 systemd services and 4 Docker containers — a cost-control playbook for indie developers.

From Streamlit to Tauri: Bringing Water Engineering Tools to the Desktop
Lessons from migrating 12 Streamlit web apps to Tauri desktop apps — separating the compute engine, rewriting the React frontend, cross-platform packaging.

Wiring 44 Scripts into Raycast
Wrapping CLI tools into one-click GUI actions — design notes, gotchas, and efficiency data from building 58 Raycast wrappers.

One-Person Full-Stack Infrastructure
24 systemd services, 4 Docker containers, 29 GitHub repos — how an indie developer builds a reliable production stack on a minimal budget.

The Zhedong Water Diversion Digital Twin: From Physical Water Network to Intelligent Scheduling
An in-depth look at Zhejiang's Zhedong Water Diversion digital twin — how digital twin technology enables precise scheduling and intelligent management of a trans-basin water network.

Welcome to My Technical Blog
My first blog post. I'll share technical experience and project insights across water conservancy, data analysis, and machine learning.
AI 方法论 · synced 10/3/2026, 20:13:44 UTC · 11 posts