Nonsense Chat GTP conversation output

63

I tried to convert ChatGPT conversations to PDF, but it's been blocked since 2026.

So I tried to build it based on the MVC model.

Since the source code is still incomplete, I created the requirements as follows.

If I just had the requirements and tried to create one file, it would go astray and it was difficult to add features.

So I roughly created the screen with the basic MVC model code I used when building other programs and applied the requirements below. It turned out better than when I used ChatGPT in the winter of 2025.

You should be able to create it by refining the requirements below. I created it by launching Chromium to generate a JSON file and outputting PDF or Word based on it.

1. GUI

1.1. GUI is created with Python Tkinter

- User enters a link and clicks a button to start the task.

- Code modification/creation proceeds only when the user clearly requests it.

1.2. Input target is ChatGPT shared link

- Basically targets shared links in the form of https://chatgpt.com/share/... format.

- No login or permission bypass methods are used.

1.3. User just needs to enter the link and click the PDF generation button

- After that, message loading, collection, and PDF generation proceed automatically.

- User does not need to scroll manually or pre-load messages.

2. Scrolling

2.1 Conversations are collected gradually while auto-scrolling

- Collect currently visible messages.

- Scroll a certain amount.

- Additionally collect newly appeared messages.

- Repeat this until the end.

- Messages combine user questions and ChatGPT answers as one message.

2.2 Wait time after scrolling is specified as a constant

- Allow users to directly change the constant value at the top of the source.

- If omission or loading problems occur, increase the wait time and re-run.

2.3 Scroll amount is set smaller than full screen

- For example, only move about 60-70% of screen height.

- Reduces message loss risk by overlapping previous and next screens partially.

3. Messages

3.1 Include message continuity verification function

- Check if the last message of the previous section exists partially in the next section.

- If there is no overlap between previous and next sections, judge as possible omission.

- Retry or leave error log if necessary.

3.2 Duplicate messages are automatically removed during collection

- If message ID is available, use ID as standard.

- If ID cannot be used, determine if it is the same message by author + message content + hash, etc.

- Logical message ID = ID managed by the program that combines user questions + corresponding ChatGPT answers.

- If the same logical message is repeatedly collected in multiple scroll sections, keep only one.

3.3 Remove duplicate content inside messages

- Prevent duplicate paragraphs, lists, tables, and code blocks from being collected within a single logical message.

- Do not simultaneously collect identical content from parent and child elements in the DOM.

- Prioritize collecting the smallest meaningful content blocks such as paragraphs, lists, tables, and code.

- If identical content is found simultaneously in parent and child blocks, keep only one based on the child block.

- If duplication is confirmed by parent/child DOM relationship, identical DOM identifiers, or duplicate elements created at the same location, keep only one.

- Do not remove different normal content just because the text is the same.

- Exclude hidden DOM, display duplicate DOM, and accessibility duplicate DOM from collection targets.

3.4 Allow section overlap during collection to prevent omission

- During auto-scroll, intentionally overlap and collect previous and next sections.

- Remove overlapped duplicate messages and content from final collected data.

- Do not leave unintended duplicates in the final collected text.

- Intentional boundary overlaps between PDFs are treated as separate normal operations according to the rules in 5.3-5.5 and 7.2

4. Images, Tables, Hyperlinks

4.1 Ignore images

- Image downloads and PDF insertions are not required.

- Delays in loading due to images, size adjustments, and page layout issues are excluded from consideration.

4.2 Text-centric PDF generation

- Targets general sentences, lists, tables, code, formulas, etc.

- Replicating the ChatGPT's special UI or interactive elements exactly as they appear on the screen is not the goal.

- Code blocks are rendered in separate code boxes distinct from the main text.

- If the code language is identified, language information such as Python or JavaScript is preserved.

- Syntax highlighting is applied to the code whenever possible.

- ChatGPT's execution buttons, copy buttons, and other interactive UI elements are not included in the PDF.

4.3 Hyperlinks only need to exist on one side of the split PDF

- Links do not need to be present in both PDFs for overlapping areas.

- Even if link text is duplicated, the actual link can remain only in the file where it first appears.

5. PDF

5.1 PDFs are generated based on A4 standards

- Paper size A4.

- Font, margins, line spacing, etc., are fixed consistently.

- The browser/Chromium rendering engine is used for final typesetting.

- Code blocks use a fixed-width font.

- Background color, borders, and internal padding are applied to code blocks.

- Syntax highlighting colors can be applied to keywords, strings, numbers, comments, etc.

5.2 The number of pages per PDF is set as a constant

- Example: 4 pages.

- Set as a constant value at the top of the source code instead of being inputted through the GUI every time.

- The specified page count is prioritized as an absolute upper limit.

- If a code block exceeds the specified page count, internal splitting of the code block is allowed.

- Only the specified number of pages are generated, and the remaining content is carried over to the next PDF generation.

- Overlapping content is also included within the specified maximum page count.

- If a single table's actual A4 rendering result exceeds the specified page count, the current implementation may not be able to maintain that table as a single block, resulting in an error where new content cannot be added to the PDF.

- Therefore, when changing the PAGES_PER_PDF value, consider the possibility that a single table, code, or formula block may exceed the specified page count.

5.3 The maximum height of overlapping areas between PDFs is set as a constant

- Example: 10mm.

- Users should be able to change the constant value at the top of the source code.

- The overlapping area does not exceed the specified height.

5.4 Intentionally overlap boundary areas when splitting PDFs

- Safe overlapping area at the end of 001.pdf

- Same area included at the beginning of 002.pdf

- The maximum height of the overlapping area follows the constant value in 5.3.

- The purpose is to prevent boundary omissions and facilitate verification.

5.5 Overlapping areas are determined based on the actual A4 rendering height

- Do not use the number of lines or characters displayed on the browser screen as a basis.

- Within the maximum overlapping height specified in 5.3, the largest safe text unit from the last part of the previous PDF is duplicated at the beginning of the next PDF.

- Sentences or structured blocks are not forcibly truncated to precisely fill the specified maximum overlapping height.

- For structured blocks such as tables, code, and formulas, the splitting rules in section 6 are prioritized.

- If a code block is split at a page boundary and continues in the next PDF, the continuing code portion is not calculated as part of the overlapping area in 5.3.

6. Page Splitting

6.1 Safe Cut Priority

- If possible:

- End of message

- End of paragraph

- End of sentence

- End of regular line

- Do not cut within structured blocks.

6.2 Tables are not cut in the middle

- If a page split point is inside a table, the entire table is moved to the next PDF if it can fit without any issues.

- Medium-sized tables (1-10 rows): Cut before the table start whenever possible at the PDF boundary and start the table in the next PDF.

- Large tables (11 rows or more): Table duplication is prohibited. Split naturally by row, and repeat only the table title/header row in the next PDF.

6.3 Code block rendering and splitting

- Code blocks are rendered as a single code box, separate from the general text.

- If the code language is detected, the language name (e.g., Python, JavaScript) is displayed at the top of the code box.

- Syntax highlighting colors are applied to the code whenever possible.

- Short or medium-length code blocks are kept as a single box as much as possible.

- If the entire code block does not fit into the remaining space of the current PDF, but it will fit in the next PDF, the entire code block is carried over to the next PDF.

- If a code block itself exceeds the maximum number of pages for a PDF, splitting at the code line level is allowed.

- Code lines are not cut in the middle as much as possible.

- Split code blocks maintain the same background color, border, monospaced font, and syntax highlighting style in the next PDF.

- The continuation of a code block in the next PDF can display the language name again or as "Python (continued)", for example.

6.4 Do not cut in the middle of formula blocks

- Formulas are treated as a single block.

- However, if a single formula block itself exceeds the maximum number of pages for a PDF, splitting is allowed while prioritizing the page limit and splitting at safe formula units or display row criteria.

6.5 Image description blocks are also kept together as much as possible

- Actual images are ignored, but if the image description text is a single block, it is kept together as much as possible.

6.6 The number of PDF pages is estimated not by the number of characters but by the actual A4 rendering result

- This is because the number of pages can vary even with the same number of characters due to tables, code, and formulas.

- The page split boundary is determined after checking the number of pages in the actual PDF rendering.

7. PDF Generation

7.1 Automatically generate multiple split PDFs

- Example:

- conversation_001.pdf

- conversation_002.pdf

- conversation_003.pdf

- …

- Each PDF is generated based on the specified number of pages.

7.2 Split PDFs include boundary overlap

- This is for verification and missing data prevention.

- The overlap is intentional and considered normal operation.

7.3 Generate a single complete PDF in the end

- A single complete PDF is also generated in the end.

- Automatic merging occurs after split PDF generation.

- Whether to keep the split PDFs is specified as a constant.

- The default setting is not to keep the split PDFs.

- Split PDFs are merged in order without modification, and boundary overlap is also maintained in the final PDF.

8. Error Handling

8.1 If an unrecoverable error occurs during conversation collection or PDF generation, the task is stopped.

- Completed PDFs and the location of the error are indicated in a message.

8.2 Display progress on the GUI

- Example:

- Current number of collected messages

- Scroll section number

- Success/failure of continuity verification

- Number of the PDF currently being generated

- Status of final merging progress

8.3 Log errors when they occur

- Message continuity failure

- Possible lack of loading wait time

- Specific PDF rendering failure

- Link access failures, etc. are logged to allow for verification.

9. Json File Requirements

- Save json files by message and scroll unit

- Message units are used for saving conversations from the middle later.

- Create conversation_title.json when finishing the work last.

- Create a temp folder under the working folder to save the message and scroll unit json files created above.

- Currently, only saving json setting files by message and scroll units is reflected.

10. Call major constants from the file

PAGES_PER_PDF = 10 # Maximum number of pages per PDF **

OVERLAP_MAX_MM = 10.0 # Maximum height (mm) to overlap with the previous PDF when splitting PDFs **

NAVIGATION_TIMEOUT_SEC = 180.0 # Initial access time limit for shared link (seconds) **

INITIAL_LOAD_WAIT_SEC = 10.0 # Wait time for chat screen loading after initial access to shared link (seconds) **

SCROLLER_SEARCH_TIMEOUT_SEC = 20.0 # Total scroller search time limit (seconds) **

SCROLL_WAIT_SEC = 10.0 # Wait time for DOM stabilization after moving to top/scrolling (seconds) If too short **

# The next scroll may be attempted before DOM height/message count changes stabilize after scrolling

# Since it takes 5 seconds to scroll through one screen,

# The scroller search time limit (SCROLLER_SEARCH_TIMEOUT_SEC) and

# The sum of SCROLL_WAIT_SEC should be less than NAVIGATION_TIMEOUT_SEC

# Previously 2 seconds, then increased to 5 seconds and then 10 seconds. **

SCROLL_RATIO = 0.4 # Scroll movement ratio relative to screen height (0.1~0.95) default 0.56 **

CONTINUITY_RETRY_COUNT = 3 # Number of retries when message continuity validation fails after scrolling **

STABLE_ROUNDS_TO_FINISH = 6 # Number of stable rounds with no changes in DOM height/message count after scrolling

MAX_SCROLL_ROUNDS = 5000 # Upper limit of scroll rounds (to prevent infinite loops)

KEEP_SPLIT_PDFS = True # Whether to keep split PDFs after final merging **

SAVE_COLLECTED_JSON = True # Whether to save collected logical messages as JSON **

VIEW_CHROMIUM = True # Whether to open Chromium DevTools for debugging (True: open, False: close)

A4_MARGIN_TOP_MM = 15.0 # PDF top margin (mm)

A4_MARGIN_BOTTOM_MM = 15.0 # PDF bottom margin (mm)

A4_MARGIN_LEFT_MM = 16.0 # PDF left margin (mm)

A4_MARGIN_RIGHT_MM = 16.0 # PDF right margin (mm)

BODY_FONT_PT = 10.5 # PDF body font size (pt)

LINE_HEIGHT = 1.55 # PDF body line height ratio

# Code block display

CODE_FONT_PT = 9.2

CODE_SYNTAX_HIGHLIGHT = True

11. Read and output Json file

11.1 Generate PDF or Word file by directly reading Json file

11.2 Create a Json checkbox above the shared link label

11.3 When Json checkbox is selected, a Find button is activated behind the shared link text box to load the Json file.

11.4 Check if the file is a usable Json file when loading.

11.5 When the PDF generation button is pressed, a PDF file is created using the json file.

11.6 When the Word generation button next to the PDF generation button is pressed, a WORD file is generated.

11.7 The WORD file button can also generate a Word file through a shared link.

12. User message

For Word and PDF, user messages are displayed in a Box etc. to distinguish them from ChatGTP messages.

99. Other

99.1 Do not print the ChatGPT page as-is with Ctrl+P

- To avoid the problem of the latter part not being printed in long conversations.

- Do not use the method of creating a PDF of the entire browser screen at once.

99.2 Direction of not creating the entire conversation as a huge HTML/PDF at once

- In very long conversations, memory/rendering issues may occur again.

- It is more stable to create independent HTML files from the collected raw data in necessary amounts and generate PDFs.

99.3 Consider the possibility of ChatGPT web structure changes

- The PDF generation part is relatively stable, but

- If ChatGPT's DOM structure or message selectors change, the collection part may need to be modified.

99.4 Current highest priority

- 1st priority: Prevent message omission

- 2nd priority: Maintain message order

- 3rd Place: Securing Duplicate Boundaries

- 4th Place: Preserving Table, Code, and Formula Blocks

- 5th Place: Accurate A4 Page Division

- 6th Place: Final PDF Merging

로그인한 회원만 댓글 등록이 가능합니다.

개발한당

KR | ID | EN
  • IDR
  • KOR
7.52 ▼ -0.01

2026.10.09 KEB 하나은행 고시회차 2027회

다가오는 한인 행사일정

  • 등록 된 일정이 없어요!