When comments scroll by, you cannot read them again
During a stream, comments arrive faster than anyone can read them. Look away for a moment to check a setting, and the message you wanted is already gone. Going back through the archive afterwards is possible, but it takes time, and what you get is often the opposite of what you wanted. The list hands you everything in the order it arrived. What you wanted was one person’s messages lined up together, not scattered between everybody else’s.
That is the small problem this article is about. I wanted to take the comment log of a stream and line up, in one place, the messages that a single person wrote one after another.
One thing I should say first. I do not run streams myself. The text I used for testing is a sample I typed by hand, and what I can report honestly is only what happened on my own desk. I cannot tell you how this behaves during a live stream, because I have not tried that.
What I tried this time: pointing the GitHub Copilot SDK at this job
The tool named in the title is the GitHub Copilot SDK. The picture in my head was simple. First, take the comment log and cut it into blocks, one block for each run of messages from the same person. Then hand those blocks to the SDK and ask it to give each block a short heading, so the log is easier to skim later.
This article covers the first half of that. In the script below, the function that builds the text for the SDK is written, but I have not yet sent that text to the SDK. So there is not a single line in this article that the SDK wrote. I would rather say that plainly than write up something I did not do. The heading step is a job for a later article, if the first half turns out to be usable.
Getting the comments into a form a script can read
Prepare a plain text file, one comment per line, in this shape:
[00:12:03] みかん: こんばんは
Three parts, separated clearly:
- the time, in square brackets, as hours, minutes and seconds
- the name of the person who wrote it
- the text of the comment
Points I checked before writing any code:
- Keep one comment on one line. If a comment itself contains a line break, join it into one line first, or the script will read only half of it.
- If a name contains a colon, the line may be split in the wrong place. In the sample I used names that contain no colons.
- Save the file as UTF-8. Names written in Japanese or other scripts will otherwise come out broken.
- Do not worry about blank lines. The script skips them.
You can make such a file by copying the log out of your own notes, or by exporting it if the tool you use offers an export. The adjusting is a little dull, but it is done once. From the second time on, you only replace the file.
The code: lining up what one person wrote in a row
This is the script I wrote. It is my own, written for this article.
"""
配信のコメントを、同じ方が続けて書いた分ごとにまとめて並べる。
GitHub Copilot SDK に渡す文も、このまとめた形から組み立てます。 SDK を入れていない手元でも、そのまま動いて結果が出ます。 使い方: python comment_blocks.py おまけのサンプルで動きます python comment_blocks.py ファイル名 手元のテキストで動きます """
import re import sys
# 1行の形:[00:12:03] みかん: こんばんは
ONE_LINE = re.compile(r"^\[(\d{2}:\d{2}:\d{2})\]\s*([^::]+?)\s*[::]\s*(.*)$")
SAMPLE_COMMENTS = """ [00:12:03] みかん: こんばんは [00:12:10] みかん: 今日は音が良いですね [00:12:22] みかん: 前は少し割れていました [00:12:35] たけのこ: わたくしも聞こえました [00:12:50] みかん: あ、たけのこさん、こんばんは [00:13:04] みかん: マイクを替えたんです [00:13:20] たけのこ: なるほど [00:13:31] みかん: 明日も同じ時間に配信します [00:13:45] かぼす: 待っています """
RULE = "─" * 34
def read_comments(text): """テキストを読み、時刻・名前・本文の組に分ける。""" comments = [] for raw in text.splitlines(): line = raw.strip() if not line: continue found = ONE_LINE.match(line) if not found: continue time, name, body = found.groups() comments.append((time, name.strip(), body.strip())) return comments
def group_by_speaker(comments): """同じ方が続けて書いた分を、1つのかたまりにまとめる。""" blocks = [] for time, name, body in comments: if blocks and blocks[-1]["name"] == name: blocks[-1]["lines"].append(body) blocks[-1]["end"] = time else: blocks.append({ "name": name, "start": time, "end": time, "lines": [body], }) return blocks
def build_prompt(blocks): """GitHub Copilot SDK に渡す文。かたまりごとに見出しを付けてもらいます。""" lines = [ "次の配信コメントを、かたまりごとに短い見出しを付けて整えてください。", "", ] for number, block in enumerate(blocks, 1): lines.append( "かたまり%d:%s(%s 〜 %s)" % (number, block["name"], block["start"], block["end"]) ) for body in block["lines"]: lines.append(" " + body) lines.append("") return "\n".join(lines)
def main(): if len(sys.argv) > 1: with open(sys.argv[1], encoding="utf-8") as handle: text = handle.read() else: text = SAMPLE_COMMENTS
comments = read_comments(text) blocks = group_by_speaker(comments)
print("まとめた配信コメント") print("読み込んだコメント:%d件" % len(comments)) print("まとめたかたまり:%d件" % len(blocks)) print()
for block in blocks: print(RULE) print("発言者:%s" % block["name"]) print("発言した時刻:%s 〜 %s" % (block["start"], block["end"])) print("まとめた発言:%d件" % len(block["lines"])) for body in block["lines"]: print(" ・%s" % body) print(RULE)
if __name__ == "__main__":
main()
Now, what each part does.
The pattern at the top. ONE_LINE is a regular expression that describes one comment line. The part for the name is written so that neither a half-width colon nor a full-width colon can be part of it, which is why the time, the name and the text split cleanly even when the writer used a full-width colon after their name. If your own log uses a different separator, this is the one line you would change.
read_comments. It walks the text line by line, skips empty lines, and skips any line that does not match the pattern. Each match becomes a set of three values: time, name, text. Nothing is guessed. A line that does not fit is dropped quietly, which is convenient, but it also means a typo in your file can make a comment vanish without any warning.
group_by_speaker. This is the heart of it. It looks at the comments in order and asks one question: is this the same name as the block just before? If yes, the text is added to that block and the end time moves forward. If no, a new block starts. Please notice the words “just before”. This does not gather everything a person said during the whole stream. It gathers only the run of messages that nobody else interrupted.
build_prompt. This builds the text that would be handed to the GitHub Copilot SDK: a short instruction at the top, then each block with its name, its start and end time, and its lines. It is not called anywhere in main, so running the script does not talk to the SDK at all.
main. If you pass a file name, it reads that file. With no file name, it uses the sample inside the script. Then it groups and prints. The printing is plain, with a rule line between blocks, because I wanted to read the result in a terminal without extra work.
What happened when I ran it
I ran the script with no file name, so it used the built-in sample of nine comments. Here is the output, exactly as it appeared.
まとめた配信コメント
読み込んだコメント:9件
まとめたかたまり:6件
──────────────────────────────────
発言者:みかん
発言した時刻:00:12:03 〜 00:12:22
まとめた発言:3件
・こんばんは
・今日は音が良いですね
・前は少し割れていました
──────────────────────────────────
発言者:たけのこ
発言した時刻:00:12:35 〜 00:12:35
まとめた発言:1件
・わたくしも聞こえました
──────────────────────────────────
発言者:みかん
発言した時刻:00:12:50 〜 00:13:04
まとめた発言:2件
・あ、たけのこさん、こんばんは
・マイクを替えたんです
──────────────────────────────────
発言者:たけのこ
発言した時刻:00:13:20 〜 00:13:20
まとめた発言:1件
・なるほど
──────────────────────────────────
発言者:みかん
発言した時刻:00:13:31 〜 00:13:31
まとめた発言:1件
・明日も同じ時間に配信します
──────────────────────────────────
発言者:かぼす
発言した時刻:00:13:45 〜 00:13:45
まとめた発言:1件
・待っています
──────────────────────────────────
The good part. Nine comments came in, six blocks came out, and that is correct. みかん wrote three times in a row at the start, and those three became one block, running from 00:12:03 to 00:12:22. The four later single messages each became a block of one, which is also right, because nobody wrote twice in a row after that. Nothing was lost and nothing was invented: three plus one plus two plus one plus one plus one is nine.
The part that did not go as hoped. Three things.
First, and most plainly, the SDK was never called. There is no AI output in this article. The function that would build the request sits in the file unused.
Second, the grouping is narrower than I first imagined. みかん appears twice in the output, once at the top and once after たけのこ’s message. If what you want is everything from one person gathered into a single place, this script does not do that. It keeps the order of the stream, and the price is that one person’s messages may be split across several blocks.
Third, the sample is small, and I typed it myself. It shows that the code runs and groups correctly. It does not show how the script behaves with a long, messy log where names repeat, where bots speak, and where a comment contains a line break. I have not tested that, so I am not going to write as if I had.
If you use this, watch these points
- The same display name is not always the same person, and one person may change their name in the middle of a stream. This script uses the name as its only clue. So it can join two people’s messages into one block, or split one person into two.
- A person who writes, waits twenty minutes, and writes again will be read as two separate blocks. If you would rather join them, you need a rule about how long a gap is still “the same run”. That rule has to be your decision, not the script’s.
- Comments from bots, notices from the streaming tool, and the text that comes attached to a donation may all sit in the same file. Decide whether to keep them before you group, or they will be mixed into somebody’s block.
- The time is only as fine as seconds. Two comments in the same second will keep the order they happen to have in the file, which may not be the true order.
- A comment text file is easy to carry around, and that is exactly the danger. It holds other people’s words, and sometimes their names. Think about whether you want to send it to a service outside your own computer, and read that service’s terms before you do.
Sources
The tool named in the title is GitHub Copilot, provided by GitHub. Its official page is https://github.com/features/copilot. The SDK for it is published by GitHub as well. I could not confirm the SDK’s own page at the time of writing, so I am not going to put a URL here that I have not checked with my own eyes. Please follow the link from GitHub’s official page rather than trusting a link I cannot stand behind. If I find the page later, I will add it to this article.
The script above is something I wrote myself for this article, and there is nobody else to credit for it. The sample comments inside it are made up, so the names みかん, たけのこ and かぼす are not real people.
Closing
I think the first half is worth keeping. A log cut into blocks is easier for me to read than a flat list, and the grouping rule is short enough to hold in your head while you look at the output. Whether the GitHub Copilot SDK then turns those blocks into useful headings, I cannot say yet, because I have not run it. When I do, I will write what happened, including the parts that did not work.
If you try this on your own comments and the script drops a line, the first thing to look at is the shape of that line. That is where the trouble has hidden for me so far.
— 辛子屋 芥子 (Karashiya Karashi)


コメント