In-Depth | Thoughts and Practices on the Application of Large Models in Data Governance AI

2023-11-13 10:00


With the rapid advancement in computer technology and the internet, data volumes have experienced exponential growth, giving rise to big data technology. The development of big data models has not only transformed the way data is processed but also altered people's thinking patterns, offering fresh perspectives and innovative approaches to problem-solving.




In this live broadcast themed "Thoughts and Practices on the Application of Large Models in AI for Data Governance," we will listen to Mr. Hu Gang as he discusses various aspects related to large models.



Data Analysis Copilot


The Copilot integrates BI with chatGPT, allowing users to simply input a statement expressing their intent to analyze specific data metrics; the model then autonomously processes and analyzes the data, quickly transforming raw data into visualized results. Similarly, the iDATA model supports creating a dashboard within minutes and is equipped for private deployment, enabling control over access and permissions, thereby ensuring data security and compliance.

Large models utilize data mining and deep learning techniques to understand business logic and datasets, generating visualizedoutputs without requiring manual code development. Enterprise users can configure component properties through a visual interface to develop business applications, resulting in an average reduction of development time by around 90%. This simplicityand speed make data analysis automation a highly efficient means to bridge the 'last mile' of automation in current practices.


Knowledge Base Construction —— Attention Is Paramount


"Attention is all you need" —— attention is the most precious capability. In today's era of information overload, individuals mustpossess highly focused attention to sieve out critical information. Corporate internal knowledge management mirrors this principle. Typically, enterprises establish centralized knowledge base platforms that collect, store, organize, and share vital information and knowledge internally.

Moreover, the application of large language models brings unique advantages to businesses. Such models demonstrate exceptional fluency in understanding and generating natural language, significantly improving accuracy compared to traditionaltechniques. They excel at grasping context and producing more coherent responses, as well as self-learning and updating based on new data. Additionally, enterprise assistants developed using large language models can assist employees in processing business data. For unstructured business data, these models can respond swiftly, greatly enhancing efficiency and accuracy in data governance. Currently, the functionalities of enterprise assistants powered by large models are gradually permeating every aspect of companies.



Application Case Study


Mr. Hu Gang illustrated through case studies the online business data processing capabilities of iDATA, also addressing issues concerning data interaction between master data systems and business systems. He pointed out that large models can serve notonly as powerful personal assistants but also provide emotional value in both professional work and personal life.

Regarding tips for querying large models, Mr. Hu offered some suggestions. Users should strive to provide clear and concise prompts, avoiding overly complex or difficult-to-understand sentences and jargon. Including contextual information in the prompt helps GPT better comprehend tasks and objectives. It's crucial to offer detailed instructions and explanations to ensure GPT can accurately complete tasks, which involves using explicit language and specific examples.



Imagining the Future


Firstly, Mr. Hu Gang quoted the lecture views of Professor Peng Kaiping, pointing out that today's children should learn to think,ask questions, and innovate, rather than simply memorize facts. With the continuous advancement of technology, we should also adjust our approach to acquiring knowledge and information, learning to utilize questioning skills and posing high-quality questions. Secondly, Mr. Hu discussed the concept of "new-quality productive forces" mentioned by General Secretary Xi Jinping, emphasizing that digital technology innovation-driven development is the main direction for forming new-quality productive forces. We should keep up with industry trends, stay informed of the latest technologies and trends, and pay attention to changes in the labor market to allocate human resources reasonably and avoid structural unemployment.

Question 1: With the introduction of large models, the personnel structure of enterprises may need to be adjusted. Mr. Hu Gang, what do you think will be the new job positions that enterprises will recruit specifically for large models in the future? Will they specifically hire prompt engineers?

Mr. Hu Gang: Prompt engineers are known as "those who know the spells." With the increasing use of large models, enterprisesare bound to increase their recruitment of prompt engineers. In addition, enterprises will also need to recruit more artificial intelligence trainers and data annotators. AI trainers need to be able to analyze product requirements and relevant data to establish data annotation rules, ultimately achieving the value of "improving the quality and efficiency of data annotation work"and "accumulating general data in specific fields." Data annotators, on the other hand, need to be able to handle various types of data, such as text, images, audio, and so on.

Among the three essential elements of AI applications - computing power, algorithms, and data - the preparation of training data can be considered the most crucial link. Andrew Ng, the founder of Google Brain, once pointed out, "80% of AI research should be focused on data preparation, ensuring data quality is the most important task. If the industry emphasizes data-centeredness rather than model-centeredness, the development of machine learning will be faster." Therefore, there will be an increased demand for personnel related to improving data quality, while low-knowledge and repetitive jobs, such as human customer service, may shrink due to the impact of large models.

Question 2: The application of large models can effectively improve work efficiency compared to traditional methods. How then should the work efficiency of employees using large models be measured?

Mr. Hu Gang: I believe we can measure the work efficiency of employees using large models from several perspectives.


①Task Completion Time: Compare the time taken by employees to complete identical tasks before and after using large models. For instance, if the use of a large model significantly reduces the time needed to complete coding tasks, then work efficiency has been improved.

②Output Quality: For natural language processing tasks, evaluate the quality of generated text or responses. If the contentproduced by the large model is more accurate and relevant, work efficiency is typically enhanced.

③Degree of Automation: Measure the extent to which an employee's workflow becomes automated when using large models. Higher levels of automation usually indicate greater efficiency since they reduce the need for manual labor.

④Error Rate: Compare error rates between using large models and traditional methods. If the use of a large model results in a lower error rate, there is potential for increased work efficiency.

⑤Customer Satisfaction: If employees interact with customers or users using large models, measure customer satisfaction through surveys or feedback. Higher satisfaction levels often correlate with higher work efficiency.

⑥Cost-effectiveness: Assess the cost-benefit of using large models versus traditional methods. If large models can provide similar or better outcomes at a lower cost, work efficiency improves. As an example, if a data project costs 2 million yuan with traditional methods, while a large model reduces it to 800,000 yuan, the cost-effectiveness is notably apparent.

⑦Task Diversity: Consider whether employees can apply large models across different types of tasks, transforming them into full-stack engineers capable of improving efficiency on a broader scale.

⑧Feedback and Improvement Speed: Evaluate how quickly employees can receive feedback and make improvements to their workflows and results when using large models.

⑨Output Volume: Consider the amount of output that an employee can produce within the same period, such as articles, reports, answers, etc














Question 3: In the PPT, it was mentioned that since China's 'population dividend' has not yet faded, the country would not preemptively invest in 'AI technology,' instead focusing more energy on research and development in areas such as new energy vehicles and chips. According to your predictions, where will China's future developments in data governanceAI technology mainly concentrate?

Mr. Hu Gang: The viewpoint expressed in the PPT reflects my personal opinion. Current chatGPT technology can lead to a significant reduction in job positions, causing the value of many ordinary workers, like CRUD Boys, SQL Boys, and news editors, to plummet dramatically. CRUD refers to the acronym for the basic operations in computing—Create, Read, Update, and Delete. CRUD is primarily used to describe fundamental operation functions within software systems or databases. While in China, chip manufacturing, new energy vehicles, and large aircraft projects can provide an abundance of job opportunities.

Regarding our future AI technology for data governance, the focus will not lie in generic AI applications like chatGPT; rather, it will be placed on proprietary large models and specific industry applications. Although the US may lead in artificial intelligence theory and groundbreaking technologies, ChatGPT, at present, is largely regarded as a 'toy,' lacking the ability to create substantial economic value. In contrast, China excels in practical AI applications, emphasizing value creation and driving continuous progress through these applications.

Huawei's Pangu large model, though not as showy as chatGPT, specializes in finance, government affairs, manufacturing, mining,meteorology, railways, and has already achieved remarkable accuracy in global weather forecasting, with instances even surpassing local manual forecasts. JD.com's YanXiao focuses on retail, finance, and supply chain logistics, NetEase Youdao's ZiYueconcentrates on education, Ctrip's WenDao specializes in tourism, and Yonyou's YonGPT is dedicated to human resource management. It is said that China currently has over 70 large-scale models under development, launch, or operation, ranging from billions to even higher parameter levels. Huawei's Pangu remains a general-purpose large model, while JD.com's YanXiao, NetEase Youdao's ZiYue, Ctrip's WenDao, and Yonyou's YonGPT are all domain-specific models with direct industrial backgrounds. They possess readily available data and application scenarios and can swiftly feedback and iterate to drive value creation—a starkly different situation from general-purpose models seeking applicable scenarios before generating value.

Moreover, specialized large models start by addressing individual problems and gradually evolve to tackle common issues, differing from general-purpose models that typically address common problems first and then adapt to specific ones. However, they might require a longer time to converge on common problems.

This distinction represents the difference between theoretical scientists and engineers. While theoretical scientists are celebratedin history, engineers are the ones who build bridges and pave roads, benefiting communities directly. Technological advancement leads to industrial upgrading, eventually translating into increased income levels, enhanced consumer confidence,and economic growth driven by internal dynamics.

Question 4: Will the introduction of large models cause employees to lose their independent ability to process data? In the future, will employees only need to know how to use models without needing to understand the underlying computational logic?

Mr. Hu Gang : The introduction of large models will not result in employees completely losing their ability to independently handle data. Instead, in the future, employees' competitive edge lies in being adept at using tools; however, to use them effectively, they must indeed grasp the underlying computational logic. Take, for example, applying Excel formulas for bid evaluation; one must already have a broad framework in mind, and then let ChatGPT fill in certain specific details. Moreover, regardless of whether it's Excel scripts or Python/Java code, these codes often contain numerous errors.

While ChatGPT possesses programming capabilities, reports have shown that its generated C++ code can be problematic, oftencontaining security vulnerabilities, and ChatGPT does not actively alert users about these issues. In knowledge base construction, ChatGPT serves as a tool that efficiently provides supporting evidence once you've already formed a well-considered perspective. The judgment and insightful wisdom still reside with the user.

ChatGPT performs well in induction and summarization, but its performance in deductive reasoning is less impressive, sometimes providing seemingly plausible but ultimately unfounded statements. During a two-day training session on Snowflake at a client's site, having studied it for barely a month myself, I found that although I could answer some questions posed by the clients, my responses were insufficiently supported and not particularly compelling. I directly consulted ChatGPT for additional arguments. ChatGPT would compile a range of evidence from various sources, and with my own judgment and logical reasoning, I could then deliver much more engaging responses to the training attendees.

When utilizing large models, such as with an AI trainer, the practitioner still needs to analyze product requirements and related data to formulate data annotation rules, ultimately aiming to enhance the quality and efficiency of data annotation work and accumulate valuable universal data in niche domains. Thus, it remains essential to understand the underlying logic and possess domain-specific knowledge or industry know-how.

Question 5: Teacher, what aspects will iDATA (iDATA) focus on in its future development?

Mr. Hu Gang: Our future focus will be on the implementation of data development Copilots and the construction of knowledgebases for industry-specific large models. More specifically, I envision a strong emphasis on the financial industry's large model sector. Currently, in the context of major B2B business domains, gpt achieves an accuracy rate of approximately 80%. We aim to elevate this accuracy rate in business domains to above 90%. In the medium term, we will work on expanding our presence in both business domains and across industries. Long-term, we will pay closer attention to the direction of intelligent agents.



Expert Introduction




Research Institutions



International Institute for Advanced Data Management Studies (IIADMS) is a non-profit, vendor-neutral institution dedicated to fostering collaboration among technology and business professionals. IIADMS is committed to advancing research in data and data management-related fields, consistently seeking new insights and best practices in the data landscape.

IIADMS is to establish itself as a preeminent global platform for knowledge exchange on theoretical and practical aspects of data management. The institution is eager to engage in diverse partnerships with prestigious domestic and international forums, both directly and indirectly addressing traditional and cutting-edge topics in data management. Through these collaborations, IIADMS aims to disseminate its research findings and contribute to the collective understanding within these forums.


The Global Data Forum 50 (GDF50) is a non-profit platform for international exchange on data management theory and practice, established under the auspices of organizations such as DAMA China. Its legal entity is authorized by the International Institute for Advanced Data Management Studies Limited, with the Forum serving as the representative responsible for the establishment, administration, and advancement of research at domestic centers.

With empowering others as its utmost objective, GDF50 regularly organizes live streaming events featuring the latest data knowledge, and has forged partnerships with governments and enterprises across China. Continuously leveraging the combined academic prowess and data-driven momentum of the Forum and the Institute, GDF50 actively contributes to the development of China's digital economy.



电话咨询:15902039750
QQ咨询:88888
微信客服
扫码咨询